Offline task scheduling method and device, electronic equipment and computer readable medium

By adopting offline task scheduling method on the deep learning algorithm platform, resource allocation is based on the priority and resource information of resource queues and resource groups, the problem of low resource utilization is solved, and more efficient resource utilization and cost reduction is achieved.

CN119938238APending Publication Date: 2025-05-06BEIJING WODONG TIANJUN INFORMATION TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202311458736.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-03
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

In the deep learning algorithm platform, due to the low utilization rate of server resources, the cost of enterprise infrastructure continues to rise. How to make full use of computing resources and improve resource utilization has become an important issue.

Method used

An offline task scheduling method is proposed. By responding to the execution of offline task resource allocation, candidate resource queues and resource groups are selected according to the priority and resource information of resource queues and resource groups, and candidate tasks are selected from them, and their resource allocation identification is set as a general identification to allocate any node resources in the cluster.

Benefits of technology

This method can improve the utilization rate of cluster resources, ensure the fairness of task resource allocation, break resource isolation, make full use of cluster idle resources, and reduce infrastructure costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938238A_ABST
    Figure CN119938238A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses an offline task scheduling method and device, electronic equipment and a computer readable medium. A specific embodiment of the method comprises the following steps: in response to execution of offline task resource allocation, selecting candidate resource queues with a first preset priority from a resource queue set according to resource information of resource queues; selecting a candidate resource group with a first preset priority from a plurality of resource groups contained in a candidate resource queue according to the resource information of the resource groups; selecting a candidate task with a first preset priority from a plurality of to-be-scheduled tasks indicated by the candidate resource group; and in response to determining that the candidate task is the preset type of task, setting a resource allocation identifier of the candidate task as a universal identifier, and allocating node resources in the cluster to the candidate task. The embodiment is related to a cluster resource scheduling technology, and any node resource can be allocated to a preset type of task by improving a task scheduling mechanism. Idle resources of the cluster can be fully utilized, and the resource utilization rate is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the technical field of cluster resource scheduling, and in particular to an offline task scheduling method, device, electronic device, and computer-readable medium. Background Art

[0002] As the industry conducts extensive research and exploration on deep learning technology, some large enterprises have developed their own one-stop algorithm platforms. Relying on cloud native technology (such as the open source container orchestration system Kubernetes), unified management of AI (Artificial Intelligence) computing power is achieved at the bottom layer, and the upper layer provides functions such as model development training, model management, and model deployment, which facilitates algorithm personnel to quickly iterate and optimize algorithms.

[0003] However, the inventors found that with the growth of training needs and the increase in the complexity of training models, the demand for the underlying computing resources of the algorithm platform is also expanding. However, due to the low utilization of server resources, it usually leads to rising corporate infrastructure costs. How to make full use of the underlying computing resources and improve the utilization of computing resources is a question of great practical significance and practical value.

[0004] The above information disclosed in this Background section is only for enhancement of understanding of the background of the inventive concept and therefore it may contain information that does not form the prior art that is already known in this country to a person of ordinary skill in the art. Summary of the invention

[0005] The content of this disclosure is used to introduce concepts in a brief form, which will be described in detail in the detailed implementation section below. The content of this disclosure is not intended to identify the key features or essential features of the technical solution claimed for protection, nor is it intended to limit the scope of the technical solution claimed for protection.

[0006] Some embodiments of the present disclosure propose an offline task scheduling method, an offline task scheduling device, an electronic device, a computer-readable medium, and a computer program product to solve one or more of the technical problems mentioned in the above background technology section.

[0007] In a first aspect, some embodiments of the present disclosure provide an offline task scheduling method, comprising: in response to executing offline task resource allocation, selecting a resource queue of a first preset priority from a resource queue set according to resource information of the resource queue as a candidate resource queue; selecting a resource group of a first preset priority from a plurality of resource groups included in the candidate resource queue according to resource information of the resource group as a candidate resource group; selecting a task to be scheduled of a first preset priority from a plurality of tasks to be scheduled indicated by the candidate resource group as a candidate task; in response to determining that the candidate task is a preset type task, setting the resource allocation identifier of the candidate task to a general identifier, and allocating node resources in a cluster to the candidate task, wherein the general identifier indicates that any node resource in the cluster can be allocated.

[0008] In some embodiments, a resource queue of a first preset priority is selected from a resource queue set based on resource information of the resource queue, including: selecting a resource queue with the highest priority from the resource queue set based on the value of the priority of the resource queue and / or the value of the resource indicator, wherein the resource indicator is used to characterize the resource usage of the resource queue.

[0009] In some embodiments, based on the priority value and / or the resource indicator value of the resource queue, a resource queue with the highest priority is selected from the resource queue set, including: based on the priority value, selecting the resource queue with the largest value from the resource queue set; in response to determining that the priority value matches, selecting the resource queue with the smallest value from the resource queue set based on the value of the first resource indicator, wherein the first resource indicator is the ratio of the used resources of the resource queue to the retained resources; in response to determining that the first resource indicator value matches, selecting the resource queue with the smallest value from the resource queue set based on the value of the second resource indicator, wherein the second resource indicator is the ratio of the used resources of the resource queue to the maximum amount of resources that can be used in excess.

[0010] In some embodiments, based on resource information of the resource group, a resource group with a first preset priority is selected from multiple resource groups included in the candidate resource queue, including: for the multiple resource groups included in the candidate resource queue, determining a third resource indicator for each resource group, wherein the third resource indicator is a ratio of the used resources of the resource group to the resource quota; and based on the value of the third resource indicator, selecting a resource group with the smallest value from the multiple resource groups.

[0011] In some embodiments, the method also includes: in response to receiving a task to be scheduled submitted by a user, dividing the data of the task to be scheduled into multiple sub-data according to the configuration parameters of the task to be scheduled; for each sub-data in the multiple sub-data, generating corresponding sub-task information according to the sub-data, and sending the sub-task information to the information queue to be scheduled for execution.

[0012] In some embodiments, the method further includes: in response to determining that the task to be scheduled is a preset type task, dividing the user into a specified resource group, and dividing the specified resource group into a specified resource queue, wherein the priority of the specified resource queue is lower than the priority of other resource queues.

[0013] In some embodiments, the method further includes: in response to determining that the allocation of idle node resources to the target task to be scheduled has failed, selecting a resource queue with a second preset priority from the resource queue set; selecting a running task from the running tasks to which node resources have been allocated under the selected resource queue as a candidate eviction task, wherein the priority of the preset type of task is lower than the priority of other tasks; based on the resource requirements of the target task to be scheduled, releasing the node resources occupied by the candidate eviction task to allocate them to the target task to be scheduled.

[0014] In some embodiments, selecting a resource queue of a second preset priority from a resource queue set includes: selecting a specified resource queue from the resource queue set, wherein the specified resource queue represents users in the resource group contained therein, who are users who submit tasks of a preset type; in response to determining that there are no running tasks under the specified resource queue, selecting an oversold resource queue with the lowest priority from the oversold resource queue set based on resource information of the resource queue, wherein the oversold resource queue is a resource queue whose used resources are greater than the reserved resources.

[0015] In some embodiments, based on the resource demand of the target task to be scheduled, releasing the node resources occupied by the candidate evicted task includes: in response to determining that the candidate evicted task is a preset type of task, for the node resources occupied by each subtask in the candidate evicted task, evicting the subtasks that match the resource demand of the target task to be scheduled; and putting the subtask information of the evicted subtask back into the information queue.

[0016] In some embodiments, based on the resource requirements of the target task to be scheduled, releasing the node resources occupied by the candidate eviction task includes: in response to determining that the candidate eviction task is another task, releasing all the node resources occupied by the candidate eviction task.

[0017] In a second aspect, some embodiments of the present disclosure provide an offline task scheduling device, comprising: a resource queue selection unit, configured to select a resource queue of a first preset priority from a resource queue set as a candidate resource queue according to resource information of the resource queue in response to executing offline task resource allocation; a resource group selection unit, configured to select a resource group of a first preset priority from a plurality of resource groups included in the candidate resource queue according to resource information of the resource group as a candidate resource group; a task selection unit, configured to select a task to be scheduled of a first preset priority from a plurality of tasks to be scheduled indicated by the candidate resource group as a candidate task; a resource allocation unit, configured to set the resource allocation identifier of the candidate task to a general identifier in response to determining that the candidate task is a preset type task, and to allocate node resources in a cluster to the candidate task, wherein the general identifier represents that any node resource in the cluster can be allocated.

[0018] In some embodiments, the resource queue selection unit is further configured to select the resource queue with the highest priority from the resource queue set according to the priority value of the resource queue and / or the value of the resource indicator, wherein the resource indicator is used to characterize the resource usage of the resource queue.

[0019] In some embodiments, the resource queue selection unit is further configured to select a resource queue with the largest value from the resource queue set based on the priority value; in response to determining that the priority value matches, select a resource queue with the smallest value from the resource queue set based on the value of a first resource indicator, wherein the first resource indicator is a ratio of the amount of used resources to the amount of retained resources of the resource queue; in response to determining that the value of the first resource indicator matches, select a resource queue with the smallest value from the resource queue set based on the value of a second resource indicator, wherein the second resource indicator is a ratio of the amount of used resources of the resource queue to the maximum amount of resources that can be used in excess.

[0020] In some embodiments, the resource group selection unit is further configured to determine a third resource indicator for each resource group included in the candidate resource queue, wherein the third resource indicator is a ratio of the used resources of the resource group to the resource quota; and based on the value of the third resource indicator, select a resource group with the smallest value from the multiple resource groups.

[0021] In some embodiments, the device also includes a task creation unit, which is configured to, in response to receiving a task to be scheduled submitted by a user, divide the data of the task to be scheduled into multiple sub-data according to the configuration parameters of the task to be scheduled; for each sub-data in the multiple sub-data, generate corresponding sub-task information based on the sub-data, and send the sub-task information to the information queue to be scheduled for execution.

[0022] In some embodiments, the device also includes a queue division unit, which is configured to divide users into designated resource groups and designated resource groups into designated resource queues in response to determining that the task to be scheduled is a preset type of task, wherein the priority of the designated resource queue is lower than the priority of other resource queues.

[0023] In some embodiments, the device also includes: a second queue selection unit, configured to select a resource queue of a second preset priority from the resource queue set in response to determining that the allocation of idle node resources to the target task to be scheduled has failed; an eviction task selection unit, configured to select a running task from the running tasks to which node resources have been allocated under the selected resource queue as a candidate eviction task, wherein the priority of the preset type of task is lower than the priority of other tasks; and a resource release unit, configured to release the node resources occupied by the candidate eviction task based on the resource requirements of the target task to be scheduled, so as to allocate them to the target task to be scheduled.

[0024] In some embodiments, the second queue selection unit is further configured to select a specified resource queue from the resource queue set, wherein the specified resource queue represents the users in the resource group contained therein, who are users who submit tasks of a preset type; in response to determining that there are no running tasks under the specified resource queue, based on the resource information of the resource queue, select the oversold resource queue with the lowest priority from the oversold resource queue set, wherein the oversold resource queue is a resource queue whose used resources are greater than the reserved resources.

[0025] In some embodiments, the resource release unit is further configured to, in response to determining that the candidate evicted task is a preset type of task, evict subtasks that match the resource requirements of the target task to be scheduled for the node resources occupied by each subtask in the candidate evicted task; and put the subtask information of the evicted subtask back into the information queue.

[0026] In some embodiments, the resource releasing unit is further configured to release all node resources occupied by the candidate eviction task in response to determining that the candidate eviction task is other tasks.

[0027] In a third aspect, some embodiments of the present disclosure provide an electronic device comprising: one or more processors; a storage device on which one or more programs are stored, and when the one or more programs are executed by one or more processors, the one or more processors implement the offline task scheduling method described in any implementation method in the above-mentioned first aspect.

[0028] In a fourth aspect, some embodiments of the present disclosure provide a computer-readable medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the offline task scheduling method described in any implementation manner in the above-mentioned first aspect is implemented.

[0029] In a fifth aspect, some embodiments of the present disclosure provide a computer program product, including a computer program, which, when executed by a processor, implements the offline task scheduling method described in any implementation manner in the above-mentioned first aspect.

[0030] The above-mentioned various embodiments of the present disclosure have the following beneficial effects: the offline task scheduling method of some embodiments of the present disclosure can improve the utilization rate of cluster resources. Specifically, in some scenarios, the relevant technology often divides exclusive resource pools for certain business groups instead of using public resource pools. For example, in Kubernetes, private node labels are marked for cluster nodes. In this way, only offline training tasks submitted by business group tenants with the node label will be scheduled to these nodes. When the computing power resources in the exclusive resource pool are idle, due to the existence of resource isolation, other business groups cannot use these idle resources, resulting in low utilization of computing power resources.

[0031] Based on this, the offline task scheduling method of some embodiments of the present disclosure can select the scheduled tasks to which resources are allocated first according to the priority of the resource queue, the priority of the resource group, and the priority of the scheduled tasks when allocating offline task resources. This can ensure the fairness of task resource allocation. For the selected scheduled tasks of a preset type, by setting the resource allocation identifier of the task to a universal identifier, any idle node resource in the cluster can be allocated to the task. In this way, when scheduling tasks, node labels can be ignored and global scheduling is supported. This can break resource isolation, make full use of cluster idle resources, and improve cluster resource utilization. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the accompanying drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that components and elements are not necessarily drawn to scale.

[0033] Figure 1 is a flow chart of some embodiments of the offline task scheduling method disclosed in the present invention;

[0034] Figure 2A These are some schematic diagrams of data splitting in the offline task scheduling method disclosed in the present invention;

[0035] Figure 2B yes Figure 1 Schematic diagrams of some application scenarios of the offline task scheduling method shown;

[0036] Figure 3 are flow charts of other embodiments of the offline task scheduling method disclosed herein;

[0037] Figure 4 yes Figure 3 Schematic diagrams of some application scenarios of the offline task scheduling method shown;

[0038] Figure 5 is an architectural diagram of an exemplary system in which some embodiments of the present disclosure may be applied;

[0039] Figure 6 It is a structural schematic diagram of some embodiments of the offline task scheduling device disclosed in the present invention;

[0040] Figure 7 It is a schematic diagram of the structure of an electronic device suitable for implementing some embodiments of the present disclosure. DETAILED DESCRIPTION

[0041] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments set forth herein. On the contrary, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not intended to limit the scope of protection of the present disclosure.

[0042] It should also be noted that, for ease of description, only the parts related to the invention are shown in the drawings. In the absence of conflict, the embodiments and features in the embodiments of the present disclosure can be combined with each other.

[0043] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0044] It should be noted that the modifications of "one" and "plurality" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, it should be understood as "one or more".

[0045] The present disclosure will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.

[0046] Figure 1 The process 100 of some embodiments of the offline task scheduling method according to the present disclosure is shown. The method comprises the following steps:

[0047] Step 101 , in response to executing offline task resource allocation, a resource queue with a first preset priority is selected from a resource queue set according to resource information of the resource queue as a candidate resource queue.

[0048] In some embodiments, the execution subject of the offline task scheduling method (such as a resource management server) can obtain the tasks to be scheduled submitted by the user through a wired connection or a wireless connection. For example, the user can submit the tasks to be scheduled to the execution subject using a terminal. For another example, the tasks to be scheduled submitted by the user can be stored on other electronic devices. The execution subject can obtain the tasks to be scheduled from the electronic device by means of message subscription. The tasks to be scheduled here are usually tasks that need to be executed and processed using cluster resources. The tasks to be scheduled can be offline tasks or other types of tasks, such as offline training tasks of models, offline reasoning and prediction tasks of models, or various other batch processing tasks, etc.

[0049] It should be noted that, at present, for offline scenarios, existing platforms are often based on the Kubernetes container orchestration system to achieve unified management of cluster resources and task scheduling. Since the Kubernetes default scheduler is only for Pod (the smallest unit that can be created and deployed in Kubernetes) level scheduling, for batch task scenarios such as AI and big data, the default scheduler Pod level scheduling usually leads to resource contention problems. That is, under limited GPU (Graphics processing unit, graphics processor) resources, a part of the GPU resources are allocated to some subtasks of task A, and another part of the GPU resources are allocated to some subtasks of task B. Since tasks A and B can only start running when all subtasks are ready, resource locks will occur. Therefore, in order to support AI batch scenarios, the existing platform further selects the open source batch scheduler Kube-Batch to support batch scheduling of offline training tasks.

[0050] In addition, for multi-tenant scenarios, existing platforms generally introduce resource group logic abstraction in the scheduling layer and resource control backend services to achieve refined management of resource quotas. In addition, for the situation where some resource groups are idle, the algorithm platform further introduces resource queue logic abstraction in the scheduling layer and resource control backend services. While ensuring the fairness of resource allocation, a resource preemption mechanism between resource queues is introduced. It supports tenants in resource queues to use resources in excess when cluster resources are idle. And when cluster resources are tight, these excess resources will be recycled to avoid idle resources. Among them, resource quotas generally specify the maximum allocable cluster resources. Cluster resources can include resources such as CPU (Central Processing Unit), memory, GPU, storage, etc. Each resource queue and resource group has a resource quota (Quota), which corresponds to the max attribute of the resource queue and the cap attribute of the resource group respectively.

[0051] In some embodiments, if the execution subject receives a task to be scheduled, or receives a subscription message indicating that there is a new task to be scheduled, or detects that the cluster has idle resources, it can be determined that the offline task resource allocation process needs to be executed. That is, cluster resources are allocated to the task to be scheduled. At this time, the execution subject can select a resource queue with a first preset priority from the resource queue set as a candidate resource queue based on the resource information of the resource queue. Among them, the resource information of the resource queue can be information that characterizes the resource situation of the resource queue, such as the resource usage (amount of used resources) and resource quota of the resource queue. The first preset priority here can be set according to actual conditions, such as the highest priority or the first priority.

[0052] As an example, the execution subject can select the resource queue with the highest priority from the resource queue set according to the priority value of the resource queue and / or the value of the resource index. Among them, the resource index can be used to characterize the resource usage of the resource queue. Which resource index to use can also be set according to the actual situation.

[0053] Specifically, when creating a resource queue, the priority of the resource queue can be set. At this time, the execution subject can select the resource queue with the largest value from the resource queue set according to the value of the priority. If it is determined that the priority values ​​match (such as the same or similar), the resource queue with the smallest value can be selected from the resource queue set according to the value of the first resource indicator. Among them, the first resource indicator can be the ratio of the used resources (used) of the resource queue to the reserved resources (min). It can be understood that in the resource queue, the resources of the reserved resources are usually guaranteed first. If it is determined that the value of the first resource indicator matches (such as the same or similar), the resource queue with the smallest value can be selected from the resource queue set according to the value of the second resource indicator. Among them, the second resource indicator can be the ratio of the used resources (used) of the resource queue to the maximum amount of resources (max) that can be used in excess. It should be noted that the overused resources are usually recycled when the resources are tight, which is manifested as the expulsion of tasks. In this way, selection according to the priority and resource indicator can ensure that tasks in resource queues with high priority or tasks in resource queues with a small proportion of resource usage are processed first.

[0054] In some optional implementations, the execution subject may also establish a priority queue containing resource information of resource queues. Then, the resource queue with the highest priority popped out of the priority queue is used as a candidate resource queue. For details, see Figure 2B The relevant description will not be repeated here.

[0055] Step 102 : Select a resource group with a first preset priority from a plurality of resource groups included in a candidate resource queue according to resource information of the resource group, as a candidate resource group.

[0056] In some embodiments, for the candidate resource queue selected in step 101, the execution subject may select a resource group with a first preset priority from multiple resource groups included in the candidate resource queue according to the resource information of the resource group, as the candidate resource group. The resource information of the resource group may be information characterizing the resource situation of the resource group, such as the resource usage (amount of used resources), resource quota, etc. of the resource group. As an example, the execution subject may establish a priority queue containing resource information of the resource group. Then, the resource group with the highest priority popped out of the priority queue is used as the candidate resource group.

[0057] Optionally, for multiple resource groups included in the candidate resource queue, the execution subject may first determine the third resource indicator of each resource group. The third resource indicator may be a ratio of the used resource amount (used) of the resource group to the resource quota (quota). Then, according to the value of the third resource indicator, the resource group with the smallest value may be selected from the multiple resource groups.

[0058] Step 103 : Select a task to be scheduled with a first preset priority from among the multiple tasks to be scheduled indicated by the candidate resource group as a candidate task.

[0059] In some embodiments, the execution subject may select a to-be-scheduled task of a first preset priority from a plurality of to-be-scheduled tasks indicated by the candidate resource group as a candidate task. That is, the execution subject may select a candidate task from the tasks to be scheduled submitted by the tenants belonging to the candidate resource group. For example, the execution subject may select the task with the highest priority from the priority queue formed by the tasks to be scheduled. As an example, the candidate tasks may be selected in the order of the amount of resources required by the tasks to be scheduled from small to large, or in the order of the time when the tasks were created.

[0060] Step 104 , in response to determining that the candidate task is a preset type task, setting the resource allocation identifier of the candidate task to a general identifier, and allocating node resources in the cluster to the candidate task.

[0061] In some embodiments, when the execution subject determines that the candidate task is a preset type task, the resource allocation identifier of the candidate task can be set as a general identifier. And the idle node resources in the cluster can be allocated to the candidate task. Among them, the general identifier can represent any idle node resources in the cluster that can be allocated. The priority of the preset type task can be lower than the priority of other tasks.

[0062] It should be noted that in order to reduce or avoid the impact on upper-level services, tasks can be graded. For some tasks to be scheduled, low-priority allocation, high-priority expulsion and global scheduling can be set, so as to make full use of idle cluster resources and improve resource utilization without upper-level services being aware of them. The preset type tasks here are also not limited and can be set according to actual conditions. For example, the preset type tasks can be mixed-deployment reasoning tasks including various offline reasoning tasks. It should be noted that the offline reasoning tasks here are not limited to model reasoning, but also applicable to various batch processing tasks. The characteristics are delay insensitivity and error re-running, such as Map Reduce (programming model for parallel computing of large-scale data sets) tasks in big data scenarios and transcoding tasks in video processing. The running results of these offline reasoning tasks usually do not cause task failures due to resource recovery, so the reliability of the results can be guaranteed.

[0063] In addition, clusters are usually shared by multiple organizations or users, which are called tenants. However, in some application scenarios, dedicated resource pools are sometimes allocated to certain business groups instead of using public resource pools. That is, the current technical solution has resource isolation. For example, in Kubernetes, resource isolation is generally achieved by marking cluster nodes with private node labels. At this time, only offline training tasks submitted by business group tenants with the node label will be scheduled to these nodes. When the computing power resources in the dedicated resource pool are idle, due to the existence of resource isolation, other business groups cannot use these idle resources, which leads to low utilization of computing power resources.

[0064] To this end, for the above-mentioned preset types of tasks, the execution entity can set the resource allocation identifier of these tasks as a universal identifier. Among them, the universal identifier representation can match the node label of any node in the cluster. In this way, the node label can be ignored when scheduling the colocation reasoning task, and global scheduling is supported, thereby breaking resource isolation, making full use of the idle resources of the cluster, and improving the utilization rate of cluster resources.

[0065] It can be seen from the above description that the offline task scheduling method in some embodiments of the present disclosure can select the scheduled tasks to which resources are preferentially allocated according to the priority of the resource queue, the priority of the resource group, and the priority of the scheduled tasks when allocating offline task resources. This can ensure the fairness of task resource allocation. For the selected scheduled tasks of a preset type, by setting the resource allocation identifier of the task to a universal identifier, any idle node resource in the cluster can be allocated to the task. In this way, when scheduling tasks, node labels can be ignored and global scheduling is supported. This can break resource isolation, make full use of cluster idle resources, and improve cluster resource utilization.

[0066] In some embodiments, in order to reduce or avoid affecting the scheduling of other tasks and achieve low priority in the allocation of preset types of tasks, when scheduling tasks, that is, the resource scheduler selects nodes with sufficient resources for the tasks, the training tasks can be given priority, followed by the mixed reasoning tasks, so as to ensure that the upper-layer business is not affected, that is, the node resources can be allocated first when the tasks are scheduled. Here, in order to facilitate the management of mixed reasoning tasks, the distinction settings can be made when creating tasks.

[0067] As an example, in response to receiving a task to be scheduled submitted by a user, the data of the task to be scheduled may first be divided into multiple sub-data according to the configuration parameters of the task to be scheduled. For example, even division may be adopted, or the execution steps in the task may be grouped according to the integrity of the task data. Then, for each sub-data in the multiple sub-data, corresponding sub-task information may be generated according to the sub-data, and the sub-task information may be sent to the information queue to be scheduled for execution. Figure 2A As shown in the figure, the tasks submitted by users 1, 2, and 3 can be split into data to obtain corresponding subtask information. In this way, the cluster can consume the subtask information and allocate cluster resources to these subtasks.

[0068] Furthermore, if it is determined that the task to be scheduled is a preset type task, the user can be divided into a specified resource group, and the specified resource group can be divided into a specified resource queue. Among them, the priority of the specified resource queue can be lower than the priority of other resource queues. In other words, the tenant who has created the mixed deployment reasoning task can be added to the dedicated resource group. And the resource group is added to the dedicated mixed deployment resource queue. And the priority attribute of the resource queue can be specified as 0. The priority attribute of the resource queue to which other tasks (such as ordinary training tasks) belong is set to 1. Therefore, when selecting resource queues during scheduling, tasks can be allocated in the order of resource queues that do not exceed the min, resource queues that exceed the min (oversold), and mixed deployment resource queues. In this way, the mixed deployment reasoning task can always be selected at the end of each round of scheduling, thereby ensuring the low-priority allocation characteristics of the mixed deployment reasoning task. It should be noted that for other tasks to be scheduled submitted by the above-mentioned users, the resource queues can be divided in a conventional manner and not placed in the mixed deployment resource queue.

[0069] As an example, in Figure 2B In the application scenario shown, in order to achieve low-priority allocation of colocation inference tasks, this solution can add a priority attribute to the resource queue, a resource allocation logic object. The specific task scheduling process is as follows:

[0070] Step 1: At the beginning of task scheduling, the execution subject can traverse all tasks to be scheduled, thereby establishing a priority queue containing resource information of resource queues and resource groups. It can be understood as a structure containing information such as resource queues and resource groups' resource usage and resource quotas. Among them, the resource information of all resource queues of a single cluster (the same cluster) can be put into the same priority queue. The resource information of all resource groups under a resource queue can be put into the same priority queue. N resource queues maintain N priority queues (heap implementation) to store their respective resource information.

[0071] Step 2: Pop the resource queue RQ with the highest priority from the priority queue RQ composed of resource queue information k (Right now ). Among them, RQ k Indicates the kth resource queue, k = 1, .., n. n indicates the total number of resource queues created in the cluster. Since the colocation resource queue has the lowest priority, it will be popped up last. If the priority queue is empty, end this round of scheduling and repeat step 1. This step solves the problem of deciding which resource queue to prioritize for tasks waiting to be scheduled when a cluster has multiple resource queues.

[0072] Step 3, Resource Queue RQ k Contains t resource groups, and the priority queue Q is formed from the resource group information corresponding to the resource queue k (Right now ) Pop up the resource group Q with the highest priority m (Right now ). Among them, Q m represents the mth resource group, m = 1,…, t. If the resource group priority queue is empty, repeat step 2. This step solves the problem of deciding which resource group to prioritize tasks submitted by when there are multiple resource groups under a resource queue. Here, in order to ensure the fairness of resource allocation, tasks in resource groups with a smaller resource occupancy ratio are prioritized.

[0073] Step 4, Resource Group Q m Contains multiple tasks to be scheduled. Select the task with the highest priority from the priority queue of tasks to be scheduled. n (i.e. Task_n). Schedule the task to the appropriate cluster node. This is task T n Each subtask Pod is bound to a node. Then, the resource queue RQ k , Resource Group Q m Put it back into the original priority queue and repeat step 2. If the task priority queue is empty, repeat step 2.

[0074] Continue to refer Figure 3, which shows a process 300 of another embodiment of the offline task scheduling method according to the present disclosure. The method may also include the following steps:

[0075] Step 301 : In response to determining that allocation of idle node resources to a target task to be scheduled fails, a resource queue of a second preset priority is selected from a set of resource queues.

[0076] In some embodiments, the execution subject of the offline task scheduling method disclosed in the present invention, when determining that the allocation of idle node resources for the target task to be scheduled has failed, can select a resource queue with a second preset priority from the resource queue set. The target task to be scheduled and the second preset priority here can also be set according to actual conditions. For example, the target task to be scheduled can be a task to be scheduled whose resource occupancy of the resource queue is within minutes (min). It should be noted that if the allocation of node resources for such a task fails, it means that there are no node resources with matching resource quantities in the cluster, or there are no idle node resources, and resources are relatively tight at this time. Therefore, task expulsion and resource preemption are required. As an example, the execution subject can select the resource queue with the smallest value (lowest priority) from the resource queue set according to the value of the priority.

[0077] Optionally, when the mixed reasoning task is stored in a separate resource queue, the execution subject can also select a specified resource queue from the resource queue set. Among them, the specified resource queue can represent the users in the resource group contained therein, and is the user who submits a preset type of task. If it is determined that there are no running tasks under the specified resource queue, that is, there are no tasks occupying resources and running. At this time, the oversold resource queue with the lowest priority can be selected from the oversold resource queue set based on the resource information of the resource queue. Among them, the oversold resource queue is generally a resource queue whose used resources are greater than the retained resources. Here, the resource indicators of the oversold resource queue can also be used for selection.

[0078] Step 302 : Select a running task from the running tasks to which node resources have been allocated under the selected resource queue as a candidate task to be evicted.

[0079] In some embodiments, based on the resource queue selected in step 301, the execution subject can select a running task from the running tasks that have been allocated node resources under the resource queue as a candidate for eviction task. For example, the candidate eviction task can be selected in the order of priority values ​​from small to large, or in the order of resources occupied by the running tasks from large to small, or in the order of the running time of the running tasks from short to long. Among them, the priority of the preset type task is usually lower than the priority of other tasks.

[0080] Step 303: based on the resource requirement of the target task to be scheduled, release the node resources occupied by the candidate eviction task to allocate them to the target task to be scheduled.

[0081] In some embodiments, based on the resource requirements of the target task to be scheduled and the attributes of the candidate task to be evicted, the node resources occupied by the candidate task to be evicted can be released, so that the released resources are allocated to the target task to be scheduled. As an example, if the candidate task to be evicted is another task, all node resources occupied by the candidate task to be evicted can be released.

[0082] In some embodiments, if the candidate evicted task is a task of the above-mentioned preset type, for the node resources occupied by each subtask in the candidate evicted task, the subtask that matches the resource demand of the target task to be scheduled can be evicted, that is, evicted on demand. And the subtask information of the evicted subtask can be put back into the information queue to be rescheduled.

[0083] The offline task scheduling method of the disclosed embodiment further enriches and improves the processing flow of task expulsion and resource release. When resources are tight, the mixed-deployment reasoning tasks that are "interrupted" can be expelled first, followed by the tasks in the oversold resource queue. This can not only ensure that resources can be released in a timely manner, but also help reduce the impact on other tasks.

[0084] It is understandable that the purpose of the hybrid reasoning task in the offline scenario is to occupy idle resources in the cluster in an "interrupted" manner to improve the utilization rate of the cluster computing resources. However, when cluster resources are tight, in order to prioritize the resource allocation of training tasks, the hybrid reasoning task needs to be able to quickly give up resources. That is, when cluster resources are tight and the resource allocation of training tasks fails, resources can be quickly given up to reduce the impact on the upper-level training business. As an example, in Figure 4 In the application scenario shown, in order to achieve high priority in the eviction of the colocation reasoning task, the specific eviction process is as follows:

[0085] Step 1: At the beginning of task eviction, the execution subject can traverse all tasks and maintain the priority queue RQ containing resource information of all oversold resource queues and co-location resource queues. Execute the resource preemption process for the training task Preemptor. Preemptor represents the task to be scheduled whose resource occupancy in the resource queue is within min. This training task is in the normal resource process ( Figure 2B (as shown in the figure) The selection of an idle node fails, so the preemption process is executed to evict the co-located inference tasks on the node and the tasks running in the oversold resource queue.

[0086] Step 2: Take out the resource queue RQ with the lowest priority from the RQ priority queue k(i.e., RQ_k). Since the priority attribute of the resource queue created for the offline colocation inference task is 0. The priority attribute value of the resource queue to which the normal training task belongs is 1. Therefore, the colocation resource queue is always selected first to ensure the high priority eviction of the colocation inference task. The purpose of this step is to determine which resource queue to evict the running tasks. The colocation inference tasks in the colocation resource queue are evicted first, followed by the oversold resource queue with a higher proportion of DRF resources, and finally the oversold resource queue with a low proportion of DRF resources. Among them, DRF is a general multi-resource maximum-minimum fairness allocation strategy that solves the problem of fair resource allocation of different types of resources in the system.

[0087] Step 3: Select the resource queue RQ from the previous step k , the running tasks submitted by all resource group tenants under this resource queue are maintained in a priority queue according to their priorities. The lowest priority task Preemptee is taken out from this priority queue. The priority here can be based on: the shorter the task starts running, or the smaller the task priority field is, the first to be evicted.

[0088] If the preemptee is a colocation inference task, the node resources occupied by the subtasks in the colocation inference task can be released as needed according to the resource requirements of the preempted task Preemptor. Figure 2A As shown, the subtask work information is put back into the information queue (message queue), so that the subtasks of the expelled colocation reasoning task can be automatically pulled up and started running when the cluster resources are idle. If the Preemptee is not a colocation reasoning task, the node resources occupied by all subtasks of the Preemptee can be released.

[0089] It is understandable that since the hybrid reasoning task will be evicted first when cluster resources are tight, the task and data shard configuration information corresponding to the hybrid reasoning task will be put back into the message queue. When cluster resources are idle, the evicted hybrid reasoning task will be restarted. The whole process is automatic and does not require user intervention, thus ensuring the correctness of the batch processing results and the reliability of the hybrid reasoning task results.

[0090] Step 4: Determine whether the resources released in step 3 meet the resource requirements of the task Preemptor. If so, end the current preemption process and execute the scheduling of the task Preemptor. If not, continue to step 2, select the next preempted task and release the resources.

[0091] See below Figure 5, which shows the architecture diagram of an exemplary system in which the offline task scheduling method of the present disclosure can be applied. The system may include Client, Task-server, Hybrid-Master, Task-status-monitor, Redis, and OSS.

[0092] Among them, Client is the command line client for submitting tasks. When submitting a hybrid inference task, specify the necessary startup parameters, such as data sharding configuration, project engineering path, runtime image name, startup script, and result output path. At the same time, the Client command line client uploads the user script to the OSS object storage.

[0093] Task-server is the server-side component for creating tasks. It creates multiple subtask messages according to the data sharding (sub-data splitting) configuration sent by the Client. Each subtask message corresponds to a data shard and is sent to the message queue.

[0094] The Hybrid-Master service is the consumer component of the message queue. The Master service takes a subtask message from the message queue and creates a corresponding Pod object for it.

[0095] The cluster scheduler listens to the creation of the Pod object and selects a node with idle resources for it. After determining the node, the node name can be filled in the nodeName field of the Pod object. The Kubelet component of the corresponding node is responsible for pulling up the Pod. When the Pod is initialized, it pulls the user-defined inference script from the OSS object storage and starts executing the inference script after the pull is completed.

[0096] Task-status-monitor is a monitoring component that monitors cluster task status information in real time based on the Informer mechanism. When the task status changes, the task status information is updated to the Redis (Remote Dictionary Server) database. For hybrid reasoning tasks, when cluster resources are tight and the task is evicted, the Monitor component sends the eviction information of the hybrid reasoning task to the Hybrid-Master service. The Master service puts the subtask information of the reasoning task back into the message queue. It will be rescheduled when cluster resources are idle.

[0097] The status information of the hybrid inference task can be stored in the Redis database. Users can use the Client command line client to obtain the status information of the offline hybrid inference task in real time.

[0098] In addition, Kubernetes generally divides dedicated node resources by setting node labels and setting Pod node affinity. That is, only when the node affinity setting of the task Pod matches the node label can the task be scheduled to the node. Therefore, in order to achieve the global scheduling characteristics of the co-location reasoning task, when creating the Pod corresponding to the co-location reasoning task, its node affinity can be set to match the node label of any node in the cluster.

[0099] It should be noted that there are the following two situations in the prior art:

[0100] First, current technical solutions generally cannot solve the phenomenon of task tidalization. For example, more offline training tasks will be submitted to the algorithm platform during the daytime hours on weekdays compared to nighttime, and on weekdays compared to non-working days. Cluster resource usage will have troughs during idle times, and cluster computing resources will be idle.

[0101] Second, the current technical solution suffers from resource fragmentation. The resource requirements for training tasks in different business scenarios vary, depending on the complexity of the model and the scale of the data input to the model training. In addition, the inconsistency in the amount of resources provided by different nodes inevitably leads to the fragmentation of cluster computing resources. For example, offline training task A applies for 3 GPU cards and task B applies for 2 GPU cards. It is known that node 1 has 4 GPU cards. When scheduling, task A is scheduled to node 1 first. At this time, node 1 has 1 GPU card remaining, which cannot be used by task B, resulting in resource fragmentation.

[0102] The offline task scheduling method of the disclosed embodiment proposes a hybrid deployment method for reasoning and training tasks in offline scenarios. The reasoning tasks here are not only applicable to model reasoning in AI scenarios, but also to batch processing tasks. When the underlying computing resources are idle, the mixed reasoning tasks are "interrupted" to run, and the occupied resources can be adaptively released quickly when the cluster resources are tight, and automatically restarted when the cluster resources are idle. The following technical effects are achieved:

[0103] First, this method supports the global scheduling of colocation inference tasks in the cluster, thereby breaking through the limitations of proprietary node labels in existing technical solutions, breaking the isolation of cluster resources, and improving resource utilization.

[0104] Secondly, this method improves the existing scheduling mechanism, and for the mixed reasoning tasks, it can allocate low-priority tasks and expel high-priority tasks, so as to achieve the upper-layer business without perception, and reduce or avoid the impact on other tasks. At the same time, when the cluster resources are idle, these idle resources can be used to schedule the mixed reasoning tasks. This can further improve the utilization rate of idle resources. It can also improve the processing efficiency of these tasks, so that they can be completed before the peak period of the task. This also helps to reduce or alleviate the task tide phenomenon in the existing technology.

[0105] Finally, this method achieves elastic scaling based on the queue placement mechanism and on-demand eviction of elastic replicas in hybrid reasoning. That is, when cluster resources are tight, the number of elastic replicas corresponding to data sharding (i.e. data splitting) is reduced. When cluster resources are idle, the number of elastic replicas is increased, so that idle resources on the cluster can be fully utilized, resource fragmentation can be reduced, and computing resource utilization can be improved.

[0106] Further references Figure 6 , as a response to the above Figures 1 to 3 The present disclosure provides some embodiments of an offline task scheduling device. Figures 1 to 3 The off-line task scheduling device can be specifically applied to various electronic devices.

[0107] like Figure 6 As shown, some embodiments of the offline task scheduling device 600 may include: a resource queue selection unit 601, configured to select a resource queue of a first preset priority from a resource queue set as a candidate resource queue according to resource information of the resource queue in response to executing offline task resource allocation; a resource group selection unit 602, configured to select a resource group of a first preset priority from a plurality of resource groups included in the candidate resource queue according to resource information of the resource group as a candidate resource group; a task selection unit 603, configured to select a task to be scheduled of a first preset priority from a plurality of tasks to be scheduled indicated by the candidate resource group as a candidate task; a resource allocation unit 604, configured to set the resource allocation identifier of the candidate task to a general identifier in response to determining that the candidate task is a preset type task, and allocate node resources in the cluster to the candidate task, wherein the general identifier indicates that any node resource in the cluster can be allocated.

[0108] In some embodiments, the resource queue selection unit 601 may be further configured to select the resource queue with the highest priority from the resource queue set according to the priority value of the resource queue and / or the value of the resource indicator, wherein the resource indicator is used to characterize the resource usage of the resource queue.

[0109] In some embodiments, the resource queue selection unit 601 can be further configured to select a resource queue with the largest value from the resource queue set based on the priority value; in response to determining that the priority value matches, select a resource queue with the smallest value from the resource queue set based on the value of the first resource indicator, wherein the first resource indicator is the ratio of the used resources of the resource queue to the retained resources; in response to determining that the value of the first resource indicator matches, select a resource queue with the smallest value from the resource queue set based on the value of the second resource indicator, wherein the second resource indicator is the ratio of the used resources of the resource queue to the maximum amount of resources that can be used in excess.

[0110] In some embodiments, the resource group selection unit 602 can be further configured to determine a third resource indicator for each resource group included in the candidate resource queue, wherein the third resource indicator is a ratio of the used resources of the resource group to the resource quota; and based on the value of the third resource indicator, select a resource group with the smallest value from the multiple resource groups.

[0111] In some embodiments, the device 600 may also include a task creation unit (not shown in the figure), which is configured to, in response to receiving a task to be scheduled submitted by a user, divide the data of the task to be scheduled into multiple sub-data according to the configuration parameters of the task to be scheduled; for each sub-data in the multiple sub-data, generate corresponding sub-task information based on the sub-data, and send the sub-task information to the information queue to be scheduled for execution.

[0112] In some embodiments, the device 600 may also include a queue division unit (not shown in the figure), which is configured to divide users into designated resource groups and divide designated resource groups into designated resource queues in response to determining that the task to be scheduled is a preset type of task, wherein the priority of the designated resource queue is lower than the priority of other resource queues.

[0113] In some embodiments, the device 600 also includes: a second queue selection unit (not shown in the figure), configured to select a resource queue of a second preset priority from the resource queue set in response to determining that the allocation of idle node resources to the target task to be scheduled has failed; an eviction task selection unit (not shown in the figure), configured to select a running task from the running tasks to which node resources have been allocated under the selected resource queue as a candidate eviction task, wherein the priority of the preset type of task is lower than the priority of other tasks; a resource release unit (not shown in the figure), configured to release the node resources occupied by the candidate eviction task based on the resource demand of the target task to be scheduled, so as to allocate them to the target task to be scheduled.

[0114] In some embodiments, the second queue selection unit can be further configured to select a specified resource queue from the resource queue set, wherein the specified resource queue represents the users in the resource group contained therein, who are users who submit tasks of a preset type; in response to determining that there are no running tasks under the specified resource queue, based on the resource information of the resource queue, select the oversold resource queue with the lowest priority from the oversold resource queue set, wherein the oversold resource queue is a resource queue whose used resources are greater than the reserved resources.

[0115] In some embodiments, the resource release unit can be further configured to, in response to determining that the candidate evicted task is a preset type of task, evict subtasks that match the resource requirements of the target task to be scheduled for the node resources occupied by each subtask in the candidate evicted task; and put the subtask information of the evicted subtask back into the information queue.

[0116] In some embodiments, the resource releasing unit may be further configured to release all node resources occupied by the candidate eviction task in response to determining that the candidate eviction task is other tasks.

[0117] It is understandable that the units recorded in the offline task scheduling device 600 are similar to those in the reference Figures 1 to 3 Therefore, the operations, features and beneficial effects described above for the method are also applicable to the device 600 and the units included therein, and will not be described in detail here.

[0118] Reference below Figure 7 , which shows a structural schematic diagram of an electronic device 700 suitable for implementing some embodiments of the present disclosure. Figure 7 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.

[0119] like Figure 7 As shown, the electronic device 700 may include a processing device 701 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded from a storage device 708 into a random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the terminal device 700 are also stored. The processing device 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.

[0120] Typically, the following devices may be connected to the I / O interface 705: input devices 706 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; output devices 707 including, for example, a speaker, a vibrator, etc.; storage devices 708 including, for example, a disk, a hard disk, etc.; and communication devices 709. The communication devices 709 may allow the electronic device 700 to communicate with other devices wirelessly or by wire to exchange data. Figure 7 The electronic device 700 is shown with various devices, but it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed instead. Figure 7Each block shown in the figure may represent one device, or may represent multiple devices as required.

[0121] In particular, according to some embodiments of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, some embodiments of the present disclosure include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In some such embodiments, the computer program can be downloaded and installed from the network through the communication device 709, or installed from the storage device 708, or installed from the ROM 702. When the computer program is executed by the processing device 701, the above-mentioned functions defined in the method of some embodiments of the present disclosure are executed.

[0122] It should be noted that the computer-readable medium recorded in some embodiments of the present disclosure may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In some embodiments of the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, device or device. In some embodiments of the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which a computer-readable program code is carried. This propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer readable signal medium may also be any computer readable medium other than a computer readable storage medium, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device. The program code contained on the computer readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0123] In some embodiments, the client and the server may communicate using any currently known or future developed network protocol such as HTTP (Hyper Text Transfer Protocol), and may be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or future developed network.

[0124] The computer-readable medium may be included in the electronic device; or it may exist independently without being assembled into the electronic device. The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device: in response to executing offline task resource allocation, select a resource queue of a first preset priority from the resource queue set according to the resource information of the resource queue as a candidate resource queue; in response to the resource information of the resource group, select a resource group of a first preset priority from a plurality of resource groups included in the candidate resource queue as a candidate resource group; select a task to be scheduled of a first preset priority from a plurality of tasks to be scheduled indicated by the candidate resource group as a candidate task; in response to determining that the candidate task is a preset type task, set the resource allocation identifier of the candidate task to a general identifier, and allocate node resources in the cluster to the candidate task, wherein the general identifier indicates that any node resource in the cluster can be allocated.

[0125] In addition, computer program code for performing the operations of some embodiments of the present disclosure may be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0126] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present disclosure. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some implementations as replacements, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0127] The units described in some embodiments of the present disclosure may be implemented by software or by hardware. The described units may also be provided in a processor, for example, may be described as: a processor including a resource queue selection unit, a resource group selection unit, a task selection unit, and a resource allocation unit. Among them, the names of these units do not constitute limitations on the units themselves in certain circumstances, for example, the resource queue selection unit may also be described as "a unit for selecting a resource queue of a first preset priority from a resource queue set".

[0128] The functions described above herein may be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.

[0129] Some embodiments of the present disclosure further provide a computer program product, including a computer program, which implements any of the above-mentioned offline task scheduling methods when executed by a processor.

[0130] The above descriptions are only some preferred embodiments of the present disclosure and an explanation of the technical principles used. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by a specific combination of the above-mentioned technical features, but should also cover other technical solutions formed by any combination of the above-mentioned technical features or their equivalent features without departing from the above-mentioned inventive concept. For example, the above-mentioned features are replaced with the technical features with similar functions disclosed in the embodiments of the present disclosure (but not limited to) and the technical solutions formed.

Claims

1. An offline task scheduling method, comprising: In response to performing offline task resource allocation, selecting a resource queue with a first preset priority from the resource queue set according to resource information of the resource queue as a candidate resource queue; Selecting, according to resource information of the resource group, a resource group with the first preset priority from a plurality of resource groups included in the candidate resource queue as a candidate resource group; Selecting a task to be scheduled with the first preset priority from a plurality of tasks to be scheduled indicated by the candidate resource group as a candidate task; In response to determining that the candidate task is a preset type task, setting the resource allocation identifier of the candidate task to a universal identifier, and allocating node resources in a cluster to the candidate task, wherein the universal identifier indicates that any node resource in the cluster can be allocated.

2. The offline task scheduling method according to claim 1, wherein: The selecting a resource queue of a first preset priority from a resource queue set according to the resource information of the resource queue comprises: A resource queue with the highest priority is selected from the resource queue set according to a priority value and / or a resource indicator value of the resource queue, wherein the resource indicator is used to characterize the resource usage of the resource queue.

3. The offline task scheduling method according to claim 2, wherein: The selecting the resource queue with the highest priority from the resource queue set according to the priority value of the resource queue and / or the value of the resource indicator includes: According to the priority value, select the resource queue with the largest value from the resource queue set; In response to determining that the priority values ​​match, selecting a resource queue with a minimum value from the resource queue set according to a value of a first resource indicator, wherein the first resource indicator is a ratio of a used resource amount to a reserved resource amount of the resource queue; In response to determining that the value of the first resource indicator matches, a resource queue with the smallest value is selected from the resource queue set according to the value of the second resource indicator, wherein the second resource indicator is the ratio of the used resources of the resource queue to the maximum amount of resources that can be used in excess.

4. The offline task scheduling method according to claim 1, wherein: The selecting, according to the resource information of the resource group, the resource group of the first preset priority from the multiple resource groups included in the candidate resource queue comprises: For the multiple resource groups included in the candidate resource queue, determine a third resource indicator for each resource group, wherein the third resource indicator is a ratio of the used resource amount to the resource quota of the resource group; According to the value of the third resource indicator, a resource group with the smallest value is selected from the multiple resource groups.

5. The offline task scheduling method according to claim 1, wherein: The method further comprises: In response to receiving a task to be scheduled submitted by a user, dividing data of the task to be scheduled into a plurality of sub-data according to configuration parameters of the task to be scheduled; For each sub-data in the plurality of sub-data, corresponding sub-task information is generated according to the sub-data, and the sub-task information is sent to an information queue to be scheduled for execution.

6. The offline task scheduling method according to claim 5, wherein: The method further comprises: In response to determining that the task to be scheduled is the preset type of task, the user is divided into a designated resource group, and the designated resource group is divided into a designated resource queue, wherein the priority of the designated resource queue is lower than the priorities of other resource queues.

7. The offline task scheduling method according to any one of claims 1 to 6, wherein: The method further comprises: In response to determining that the allocation of idle node resources to the target task to be scheduled fails, selecting a resource queue of a second preset priority from the resource queue set; Selecting a running task from the running tasks to which node resources have been allocated under the selected resource queue as a candidate for eviction task, wherein the priority of the preset type of task is lower than the priority of other tasks; Based on the resource requirement of the target task to be scheduled, the node resources occupied by the candidate eviction task are released to be allocated to the target task to be scheduled.

8. The offline task scheduling method according to claim 7, wherein: The selecting a resource queue of a second preset priority from the resource queue set includes: Selecting a designated resource queue from the resource queue set, wherein the designated resource queue represents a user in a resource group contained therein, who is a user who submits a task of the preset type; In response to determining that there is no running task under the designated resource queue, an oversold resource queue with the lowest priority is selected from the oversold resource queue set according to resource information of the resource queue, wherein the oversold resource queue is a resource queue whose used resources are greater than the reserved resources.

9. The offline task scheduling method according to claim 7, wherein: The releasing the node resources occupied by the candidate eviction task based on the resource demand of the target task to be scheduled includes: In response to determining that the candidate evicted task is the preset type of task, for the node resources occupied by each subtask in the candidate evicted task, evicting the subtask that matches the resource demand of the target task to be scheduled; and Put the subtask information of the evicted subtask back into the information queue.

10. The offline task scheduling method according to claim 7, wherein: The releasing the node resources occupied by the candidate eviction task based on the resource demand of the target task to be scheduled includes: In response to determining that the candidate eviction task is another task, all node resources occupied by the candidate eviction task are released.

11. An offline task scheduling device, comprising: A resource queue selection unit is configured to select a resource queue of a first preset priority from a resource queue set as a candidate resource queue according to resource information of the resource queue in response to performing offline task resource allocation; a resource group selection unit configured to select, according to resource information of the resource group, a resource group of the first preset priority from a plurality of resource groups included in the candidate resource queue as a candidate resource group; A task selection unit is configured to select a to-be-scheduled task of the first preset priority from a plurality of to-be-scheduled tasks indicated by the candidate resource group as a candidate task; The resource allocation unit is configured to, in response to determining that the candidate task is a preset type task, set the resource allocation identifier of the candidate task to a universal identifier, and allocate node resources in the cluster to the candidate task, wherein the universal identifier represents that any node resource in the cluster can be allocated.

12. An electronic device comprising: one or more processors; a storage device having one or more programs stored thereon, When the one or more programs are executed by the one or more processors, the one or more processors implement the offline task scheduling method as described in any one of claims 1-10.

13. A computer readable medium having a computer program stored thereon, wherein: When the computer program is executed by a processor, the offline task scheduling method as described in any one of claims 1 to 10 is implemented.

14. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements the offline task scheduling method according to any one of claims 1 to 10.

Citation Information

Cited By

  • Resource defragmentation method, controller, control node and computer cluster

    CN120950224A