Methods, apparatus, electronic devices, and media for scheduling accelerator resources

By dynamically scheduling accelerator resources in a cloud service model, the problem of accelerator resource fragmentation is solved, resource utilization is improved, waiting time for model training or inference is reduced, and costs are lowered.

CN117472570BActive Publication Date: 2026-05-26BEIJING VOLCANO ENGINE TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING VOLCANO ENGINE TECH CO LTD
Filing Date
2023-10-27
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

The low utilization rate of accelerator resources in the existing cloud service model leads to frequent fragmentation, increasing the latency and operating cost of model training or inference. Traditional improvement methods have failed to effectively solve the problem of fully utilizing accelerator resources.

Method used

By acquiring accelerator resource allocation requests from machine learning models, rescheduling conditions are triggered. Based on the strategy, container groups to be rescheduled are determined and migrated from their original nodes to another node. Resources are allocated to accelerator resource allocation requests, making full use of idle fragmented resources.

Benefits of technology

It improves the utilization efficiency of accelerator resources, reduces the waiting time for model training or inference, lowers operating costs, and increases the speed of the overall task process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117472570B_ABST
    Figure CN117472570B_ABST
Patent Text Reader

Abstract

Embodiments of this disclosure relate to methods, apparatus, electronic devices, and media for scheduling accelerator resources. The method includes acquiring an accelerator resource allocation request related to a machine learning model, wherein the accelerator resource allocation request indicates the amount of accelerator resources required by a group of containers running the machine learning model. The method further includes determining a group of containers to be rescheduled according to a rescheduling policy in response to a rescheduling condition triggered by the accelerator resource allocation request. Furthermore, the method includes allocating accelerator resources on a first node for the accelerator resource allocation request by migrating the group of containers to be rescheduled from a first node to a second node. According to embodiments of this disclosure, when an accelerator resource allocation request triggers a rescheduling condition, determining the group of containers to be rescheduled according to a rescheduling policy, and then migrating the group of containers to be rescheduled to another node to satisfy the resource allocation request, can improve the utilization efficiency of accelerator resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure generally relates to the field of computers, and more specifically to methods, apparatus, electronic devices, and media for scheduling accelerator resources. Background Technology

[0002] A machine learning model is a model that learns from large amounts of data to predict new data. It can automatically extract useful features from data and make predictions based on those features. The learning process of a machine learning model typically includes two phases: training and inference. During training, the model learns and adjusts its parameters based on known data. During inference, the model uses the known results to infer data, evaluating its accuracy and generalization ability. Machine learning models have a wide range of applications, including but not limited to image recognition, speech recognition, natural language processing, and recommender systems. For example, in image recognition, a machine learning model can learn from large amounts of image data, automatically extract features from images, and predict the category of new images. In natural language processing, a machine learning model can learn from large amounts of text data, automatically extract linguistic features from text, and predict the topic of new text.

[0003] Machine learning models have brought about a wider range of applications and higher performance capabilities, but they also have longer training and inference times, larger model sizes, and higher storage costs. Furthermore, machine learning models require greater computing power, necessitating the use of more powerful computers and computing resources, and thus incurring higher costs. Summary of the Invention

[0004] Embodiments of this disclosure provide a method, apparatus, electronic device, and medium for scheduling accelerator resources.

[0005] According to a first aspect of the disclosure, a method for scheduling accelerator resources is provided. The method includes obtaining an accelerator resource allocation request related to a machine learning model, the accelerator resource allocation request indicating the amount of accelerator resources required by a container group running the machine learning model. The method also includes determining a container group to be rescheduled according to a rescheduling policy in response to a rescheduling condition triggered by the accelerator resource allocation request. Furthermore, the method includes allocating accelerator resources on the first node for the accelerator resource allocation request by migrating the container group to be rescheduled from a first node to a second node.

[0006] In a second aspect of the disclosure, an apparatus for scheduling accelerator resources is provided. The apparatus includes a request acquisition module configured to acquire an accelerator resource allocation request related to a machine learning model, the accelerator resource allocation request indicating the amount of accelerator resources required by a container group running the machine learning model. The apparatus also includes a container group determination module configured to determine a container group to be rescheduled according to a rescheduling policy in response to a rescheduling condition triggered by the accelerator resource allocation request. Furthermore, the apparatus includes a container group migration module configured to allocate accelerator resources on a first node for the accelerator resource allocation request by migrating the container group to be rescheduled from a first node to a second node.

[0007] In a third aspect of this disclosure, an electronic device is provided. The electronic device includes a processor and a memory coupled to the processor, the memory having instructions stored therein, which, when executed by the processor, cause the electronic device to perform the method according to the first aspect.

[0008] In a fourth aspect of this disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores computer-executable instructions, which are executed by a processor to implement the method according to the first aspect.

[0009] The summary section is intended to present the chosen concepts in a simplified form, which will be further described in the detailed description below. The summary section is not intended to identify key or principal features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. Attached Figure Description

[0010] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:

[0011] Figure 1 A schematic diagram of an example environment in which some embodiments of this disclosure may be implemented is shown;

[0012] Figure 2 Flowcharts of methods for scheduling accelerator resources according to some embodiments of this disclosure are shown;

[0013] Figure 3 A schematic diagram illustrating the allocation of accelerator resources according to some embodiments of this disclosure is shown;

[0014] Figure 4 A schematic diagram illustrating the workflow of a method for scheduling accelerator resources according to some embodiments of this disclosure is shown;

[0015] Figure 5A schematic diagram of the architecture of an example migration container group is shown, representing some embodiments of this disclosure.

[0016] Figure 6 The diagram illustrates the interaction between the scheduler and the daemon within a container group, representing some embodiments of this disclosure.

[0017] Figure 7 Block diagrams of apparatus for scheduling accelerator resources according to some embodiments of the present disclosure are shown; and

[0018] Figure 8 A block diagram of an example electronic device according to some embodiments of the present disclosure is shown.

[0019] In all the accompanying figures, the same or similar reference numerals denote the same or similar elements. Detailed Implementation

[0020] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.

[0021] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.

[0022] For example, upon receiving a user's proactive request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.

[0023] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0024] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.

[0025] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0026] In the description of embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The terms "first", "second", etc., may refer to different or the same objects unless explicitly stated. Other explicit and implicit definitions may also be included below.

[0027] To enable more accelerator resources to participate in machine learning model development, various service models have emerged. Infrastructure as a Service (IaaS) provides IT infrastructure as a service over a network, allowing users to deploy and run any software, including operating systems and applications, on a paid basis. Platform as a Service (PaaS) provides server platforms as a service. In PaaS, users do not need to manage or control the underlying infrastructure, including networks, servers, operating systems, and storage, but they can control the deployed applications and potentially the configuration of the hosting environment running those applications. PaaS and similar service models can employ finer-grained resource management methods, such as container groups or container-level fine-grained management. In container-level resource management, users only need to focus on the business logic within the container group or container; that is, users only need to focus on their data or models, training quality, and inference quality. In IaaS, once a user has paid, the cloud service provider will not interfere with the user's use of resources, regardless of whether the infrastructure is idle or under load. In the Model-as-a-Service (MaaS) framework, cloud service providers only benefit when users utilize the relevant resources. If there are a large amount of idle or fragmented computing resources on the platform, the cloud service provider's resource utilization will decrease.

[0028] Traditional cloud service models often fail to fully utilize accelerator resources in their algorithms, leading to frequent fragmentation. This not only increases latency during model training and inference but also wastes computing power and increases operational costs. One improvement involves unlimited accelerator resource allocation to ensure the cluster continues to meet computing demands, but this doesn't solve the fragmentation problem, and the high cost of accelerator resources adds further expense. Another approach is accelerator resource virtualization, but this doesn't dynamically balance resource usage across the entire cluster. A third approach involves manual intervention in accelerator resource selection, but this incurs additional overhead. Therefore, to avoid the problem of abundant idle fragmented accelerator resources that could meet computing demands but remain underutilized, a method is needed to automatically resolve accelerator resource fragmentation within the cluster.

[0029] To this end, embodiments of this disclosure, during the scheduling of accelerator resources, acquire accelerator resource allocation requests related to machine learning models. These requests indicate the amount of accelerator resources required by the container group running the machine learning model. If the accelerator allocation request triggers rescheduling conditions, the container group to be rescheduled is determined according to a rescheduling strategy. Then, the container group to be rescheduled is migrated from its original node to another node, thereby allocating accelerator resources to the container group running the machine learning model at the original node. Therefore, embodiments of this disclosure can automatically and fully utilize idle fragmented accelerator resources, effectively accelerating the execution of machine learning model training or inference tasks and improving the utilization efficiency of accelerator resources.

[0030] Figure 1 A schematic diagram of an example environment 100 in which some embodiments of this disclosure may be implemented is shown. The example environment of this disclosure includes at least multiple accelerator clusters and a scheduler, wherein several accelerators are configured in the accelerator clusters. In some embodiments of this disclosure, a graphics processing unit (GPU) is used as an example of an accelerator resource; however, embodiments of this disclosure may also be used in combination with other accelerator resources. Reference Figure 1 Example environment 100 includes at least GPU cluster 110, GPU cluster 120, GPU cluster 130, and scheduler 140.

[0031] like Figure 1As shown, GPU cluster 110 includes at least GPU 110-1, GPU 110-2, GPU 110-3, GPU 110-4, GPU 110-5, GPU 110-6, GPU 110-7, and GPU 110-8. It should be understood that a GPU cluster may include more or fewer GPUs. GPU cluster 120 includes at least GPU 120-1, GPU 120-2, GPU 120-3, GPU 120-4, GPU 120-5, GPU 120-6, GPU 120-7, and GPU 120-8. GPU cluster 130 includes at least GPU 130-1, GPU 130-2, GPU 130-3, GPU 130-4, GPU 130-5, GPU 130-6, GPU 130-7, and GPU 130-8. In some embodiments, the GPUs in GPU cluster 110, GPU cluster 120, or GPU cluster 130 are homogeneous, i.e., GPUs of the same type. Alternatively, the GPUs in GPU cluster 110, GPU cluster 120, or GPU cluster 130 are heterogeneous, i.e., the GPU cluster includes multiple GPUs of different types or from different manufacturers.

[0032] In some embodiments, GPU clusters 110, 120, and 130 are configured on different nodes, where a single container or container group can use one or more GPUs within the same node. In a GPU cluster, a container group is the smallest basic unit of deployment and management in the cluster, containing one or more containers. Normally, once a container group is bound to a node, it is not rescheduled. Machine learning model files are loaded into the GPU cluster for computation. For example, a container can utilize five GPUs in GPU cluster 110. These five GPUs could be GPU 110-1, GPU 110-2, GPU 110-3, GPU 110-4, and GPU 110-5, or they could be GPU 110-5, GPU 110-6, GPU 110-7, and GPU 110-8.

[0033] In some embodiments, when allocating GPU resources, the scheduler prioritizes allocating GPU resources at the entire machine without restarting the entire machine's accelerator resources. For example, if GPUs in GPU clusters 110, 120, and 130 are all idle, and a container group requests the allocation of 3 GPU resources, the scheduler can allocate the 3 idle GPUs in GPU cluster 110 to that container group. Furthermore, if another container group requests the allocation of 1 GPU resource, the scheduler will allocate one of the 5 idle GPUs in GPU cluster 110 to that container group, instead of allocating GPU resources from GPU clusters 120 and 130.

[0034] In some embodiments, the entire machine learning task is configured on the same node. For example, accelerator cluster 110 has 3 idle GPUs, accelerator cluster 120 has 2 idle GPUs, and accelerator cluster 130 has 1 idle GPU. When a machine learning task request is received that requires the allocation of 4 GPU resources, the task of the container group occupying 3 GPU resources in GPU cluster 130 is migrated to the 3 idle GPUs in GPU cluster 110, and then the machine learning task container group is configured in GPU cluster 130. This avoids the latency caused by cross-node communication between GPUs and makes full use of the fragmented GPUs in the GPU cluster. In some embodiments, model files used for inference or training are scheduled by the scheduler to the GPU cluster for computation or to perform other tasks.

[0035] It should be understood that the architecture and functionality in example environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of this disclosure. Embodiments of this disclosure can also be applied to other environments with different structures and / or functionalities.

[0036] The following will combine Figures 2 to 8 The process according to embodiments of this disclosure is described in detail. For ease of understanding, the specific data mentioned in the following description are exemplary and not intended to limit the scope of this disclosure. It is understood that the embodiments described below may also include additional actions not shown and / or actions shown may be omitted, and the scope of this disclosure is not limited in this respect.

[0037] Figure 2 A flowchart of a method 200 for scheduling accelerator resources according to some embodiments of this disclosure is shown. The method can be... Figure 1 The scheduler 140 described is used for execution. In box 202, an accelerator resource allocation request related to the machine learning model is obtained, wherein the accelerator resource allocation request indicates the amount of accelerator resources required by the container group running the machine learning model. For example, refer to... Figure 3If a container group running a machine learning model requires 16 GPU resources, then request 351 to allocate 16 GPU resources to the accelerator resource set. In some embodiments, accelerator resource allocation requests related to the machine learning model can be obtained from the scheduler.

[0038] In box 204, in response to an accelerator resource allocation request triggering a rescheduling condition, the container group to be rescheduled is determined according to the rescheduling policy. In some embodiments, the determination of whether a rescheduling condition has been triggered can be based on the relationship between the accelerator resource allocation request and the number of accelerator resources in the accelerator set. For example, refer to... Figure 3 Accelerator resource sets 310, 320, 330, and 340 each contain 8 accelerator resources, with 6 currently available. Upon receiving accelerator resource request 356 requesting the allocation of 4 GPU resources, it is determined whether request 356 triggers the rescheduling condition. For example, if the number of GPUs required by request 356 (4) is less than the number of available accelerator resources (6) in the accelerator sets, and the number of GPUs required by request 356 (4) is greater than the remaining resources configured in each accelerator resource set, then request 356 has triggered the rescheduling condition.

[0039] In some embodiments, the container group to be rescheduled is determined based on the relationship between the accelerator resources required by the accelerator resource allocation request and the number of accelerator resources in the accelerator set. (Continue to refer to...) Figure 3 At this time, the number of idle accelerator resources in accelerator set 320 is 3, which is greater than or equal to the number of accelerator resources occupied by the container group identified as to be scheduled in accelerator set 340, which is 3. Furthermore, the number of idle accelerator resources in accelerator set 340 after releasing 3 accelerator resources is 6, which is greater than or equal to the number of accelerator resources required by accelerator request 356, which is 4. Therefore, the container group occupying 3 accelerator resources in accelerator resource set 340 is determined to be the container group to be scheduled, and the node where accelerator set 320 is located is determined to be the target node of the container group to be rescheduled.

[0040] In some embodiments, if there are multiple container groups to be rescheduled, the container group with the shortest time can be selected as the container group to be rescheduled. For example, see reference. Figure 3When accelerator resource allocation request 356 triggers the rescheduling condition, the scheduling policy can select either the container group occupying 2 accelerator resources in accelerator set 340 as the container group to be rescheduled, or the container group occupying 3 accelerator resources in accelerator set 340 as the container group to be rescheduled. If the time taken to reschedule the container group occupying 2 accelerator resources in accelerator set 340 to the target node is shorter than the time taken to reschedule the container group occupying 3 accelerator resources in accelerator set 340 to the target node, then the container group occupying 2 accelerator resources in accelerator set 340 is determined to be the container group to be rescheduled.

[0041] In some embodiments, if multiple container groups are to be rescheduled, the container group that will not affect the system task process can be selected as the rescheduled container group. For example, the container group executing training tasks can be selected as the rescheduled container group, or the container group executing backend tasks can be selected as the rescheduled container group, thus not affecting normal network service provisioning. In some embodiments, if multiple container groups to be rescheduled are all container groups executing inference tasks, the container group with lower priority is selected as the rescheduled container group. In some embodiments, the container group used for hosting network services or stateless services can be directly stopped, and then the container group can be quickly restarted on another node.

[0042] In box 206, accelerator resources are allocated on the first node for accelerator resource allocation requests by migrating the container group to be rescheduled from the first node to the second node. (See reference) Figure 3 Accelerator resource request 356 requests the allocation of 4 GPU resources. The container group in accelerator set 340 that occupies 3 GPU resources is scheduled from the node where accelerator set 340 is located to the node where accelerator set 320 is located. Then, on the node where accelerator set 340 is located, the GPU resources are allocated to accelerator resource request 356. According to the scheme of this disclosure, the utilization rate of accelerator resources can be improved.

[0043] In some embodiments, before migrating the container group to be rescheduled from the original node to the target node, it is also necessary to stop the processes within the container group to be rescheduled and save the state of the relevant runtime data. For example, refer to Figure 5 When migrating container group 512 from host 570 to host 580, it is necessary to stop MaaS 514 within the container group and save its runtime data status, such as checkpoints, to the shared file system 590.

[0044] In some embodiments, the process of stopping the interaction between the scheduler and the daemons within the container group is stopped for a group of containers to be rescheduled. For example, see [reference]. Figure 6The scheduler 620 sends a stop command to the container group 616. Upon receiving the command, the daemon 617 within the container group 616 saves the status file of the running data within the container group to the file system and then notifies the scheduler that the above operation has been completed. In some embodiments, the status file of the running data refers to a checkpoint.

[0045] In some embodiments, after allocating resources to the container group to be rescheduled and allocating resources to the accelerator resource allocation request, the container group to be rescheduled and the container group running machine learning model files can be started simultaneously. In some embodiments, the container group to be rescheduled and the container group running machine learning model files have equal resource request priority, and this priority is higher than the resource request of any other container group. According to the scheme of this disclosure, fragmented accelerator resources in the cluster can be effectively utilized, waiting time can be reduced, thereby speeding up the overall task process.

[0046] Figure 3 A schematic diagram illustrating the allocation of accelerator resources 300 according to some embodiments of this disclosure is shown. For example... Figure 3 As shown, each of the accelerator resource sets 310, 320, 330 and 340 is equipped with 8 accelerators.

[0047] refer to Figure 3 Accelerator resource allocation request 351 requests the allocation of 16 accelerator resources. At this time, the accelerators in accelerator resource sets 310, 320, 330, and 340 are all idle, and the idle accelerator resources in accelerator resource sets 310 and 320 can be allocated to request 351. After allocation 361, accelerator resource set 310 has 0 idle accelerators, accelerator resource set 320 has 0 idle accelerators, accelerator resource set 330 has 8 idle accelerators, and accelerator resource set 340 has 8 idle accelerators.

[0048] Next, the next accelerator resource allocation request 352 requests the allocation of 6 accelerator resources. At this time, accelerator resource set 310 has 0 idle accelerators, accelerator resource set 320 has 0 idle accelerators, accelerator resource set 330 has 8 idle accelerators, and accelerator resource set 340 has 8 idle accelerators. The idle accelerator resources in accelerator resource set 330 can be allocated to request 352. After allocation 362, accelerator resource set 310 has 0 idle accelerators, accelerator resource set 320 has 0 idle accelerators, accelerator resource set 330 has 2 idle accelerators, and accelerator resource set 340 has 8 idle accelerators.

[0049] Continue to refer to Figure 3 Accelerator resource allocation request 353 requests the allocation of 4 accelerator resources. At this time, accelerator resource set 310 has 0 idle accelerators, accelerator resource set 320 has 0 idle accelerators, accelerator resource set 330 has 2 idle accelerators, and accelerator resource set 340 has 8 idle accelerators. The idle accelerator resources in accelerator resource set 340 can be allocated to request 353. After allocation 363, accelerator resource set 310 has 0 idle accelerators, accelerator resource set 320 has 0 idle accelerators, accelerator resource set 330 has 2 idle accelerators, and accelerator resource set 340 has 4 idle accelerators.

[0050] Next, refer to Figure 3 A new accelerator resource allocation request 354 requests the allocation of 3 accelerator resources. At this time, accelerator resource set 310 has 0 idle accelerators, accelerator resource set 320 has 0 idle accelerators, accelerator resource set 330 has 2 idle accelerators, and accelerator resource set 340 has 4 idle accelerators. The idle accelerator resources in accelerator resource set 340 can be allocated to request 354. After allocation 364, accelerator resource set 310 has 0 idle accelerators, accelerator resource set 320 has 0 idle accelerators, accelerator resource set 330 has 2 idle accelerators, and accelerator resource set 340 has 1 idle accelerator.

[0051] At this time, Figure 3 In the example, accelerator resource release request 355 requests accelerator resource set 320 to release 3 accelerator resources. After release 365, at this time, accelerator resource set 310 has 0 idle accelerators, accelerator resource set 320 has 3 idle accelerators, accelerator resource set 330 has 2 idle accelerators, and accelerator resource set 340 has 1 idle accelerator.

[0052] Next, a new accelerator resource allocation request 356 requests the allocation of 4 accelerator resources. At this point, accelerator resource set 310 has 0 idle accelerators, accelerator resource set 320 has 3 idle accelerators, accelerator resource set 330 has 2 idle accelerators, and accelerator resource set 340 has 1 idle accelerator. Although there are a total of 6 accelerator resources in the accelerator resource sets, the idle accelerator resources are distributed across 3 different accelerator sets. To use these distributed accelerator resources, remote communication needs to be established between these accelerators, such as through Transmission Control Protocol / Internet Protocol (TCP / IP), Remote Direct Data Access (RDMA), or directly using the communication channels between GPUs. This approach will degrade the performance of the machine learning model because remote communication increases communication latency.

[0053] According to an embodiment of this disclosure, processes previously allocated to three accelerator resources in accelerator resource set 340 can be scheduled to accelerator resource set 320, and then accelerators in accelerator resource set 340 can be allocated to request 356. After rescheduling (366), accelerator resource set 310 has 0 idle accelerators, accelerator resource set 320 has 0 idle accelerators, accelerator resource set 330 has 2 idle accelerators, and accelerator resource set 340 has 4 idle accelerators. Then, the 4 idle accelerators in accelerator resource set 340 are allocated to request 356. After allocation (367), accelerator resource set 310 has 0 idle accelerators, accelerator resource set 320 has 0 idle accelerators, accelerator resource set 330 has 2 idle accelerators, and accelerator resource set 340 has 0 idle accelerators. In this way, fragmented accelerator resources in the cluster can be fully utilized, thereby maximizing the utilization rate of accelerator resources without reducing the user experience of using accelerator resources.

[0054] In some embodiments, before scheduling a container group from accelerator resource set 320 to accelerator resource set 330, it is necessary to save the container group's running data state, such as checkpoints, to the file system and stop the container group from running. In some embodiments, the container group to be rescheduled has the same resource request priority as the new container group, and this priority is higher than the resource request priority of other container groups besides the container group to be rescheduled and the new container group.

[0055] Figure 4 A schematic diagram illustrating the workflow of a method 400 for scheduling accelerator resources according to some embodiments of this disclosure is shown. At block 402, the process of dynamically allocating accelerator resources is initiated. At block 404, an accelerator resource allocation request is requested, obtaining an accelerator resource allocation request related to a machine learning model file, the request representing the amount of accelerator resources required to run the machine learning model file. See, for example, [reference needed]. Figure 3 Accelerator resource allocation request 351 requests allocation of 16 GPU resources.

[0056] In box 406, determine whether there are any full-system accelerator resources available to satisfy the obtained accelerator resource allocation request. For example, if a request is made to allocate 6 accelerator resources, and there are 8 idle accelerators in accelerator set 330 that can satisfy the request, then the result is yes. As another example, if a request is made to allocate 4 accelerator resources, and there are only 3 idle accelerators in the accelerator set that can satisfy the request, then the result is no.

[0057] If it is determined in box 406 that there are available full-system accelerator resources to satisfy the obtained accelerator resource allocation request, then in box 410, the accelerator resources are directly allocated. For example, if a request is made to allocate 6 accelerator resources, and there are 8 idle accelerators in accelerator set 330 that can satisfy the request, then the 6 idle accelerators in accelerator resource set 330 are allocated to the request. In box 416, the accelerator resources are used to process the requested task. For example, if the accelerator resource allocation request is for the allocation of 4 accelerator resources, and there are 8 idle accelerators in accelerator set 330 that can satisfy the request, then in box 410, 4 idle accelerator resources in accelerator set 330 are allocated to the request, and then in box 416, the allocated accelerator resources are used to process the task request. For example, when requesting the allocation of 3 accelerator resources, if there are 4 idle accelerators in accelerator set 330 that can satisfy the request, then in box 410, 4 of the idle accelerator resources in accelerator set 330 are allocated to the request, and then in box 416, the allocated accelerator resources are used to process the task request.

[0058] If it is determined in box 406 that there are no full-system accelerator resources that satisfy the obtained accelerator resource allocation request, then in box 408, it is determined whether there are fragment accelerator resources that satisfy the request. If not, then in box 414, accelerator resources are added. If they exist, then in box 412, a container group to be rescheduled is searched. In some embodiments, a request to allocate accelerator resources is obtained in box 404, for example, a new container group requests the allocation of 8 accelerator resources. Then, in box 406, it is determined whether there are full-machine accelerator resources in the accelerator sets that can satisfy the request. For example, if there are no full-machine accelerator resources in accelerator sets 310, 320, 330, and 340 that can satisfy the request, then in box 408, it is determined whether there are fragmented accelerator resources that can satisfy the request. For example, if there are 6 remaining accelerator resources in accelerator sets 310, 320, 330, and 340 that cannot satisfy the request, and after a certain number of searches (e.g., 3 times), not enough fragmented accelerator resources are released to satisfy the request of the new container group, then in box 414, accelerator resources are added, for example, adding an accelerator cluster configured with 8 accelerator resources. Then, in box 410, the accelerator resources are allocated to the request, for example, allocating the added 8 accelerator resources to the new container group request. Through this accelerator resource allocation method, the fragmented accelerator resources in the cluster can be fully utilized without waiting for the release of other accelerator resources, thereby maximizing the utilization rate of accelerator resources.

[0059] In some embodiments, when request 356 requests the allocation of 4 accelerator resources, and there are 6 remaining accelerator resources in accelerator sets 310, 320, 330, and 340 that can satisfy the request, then in box 412, a container group to be rescheduled is searched. For example, a container group occupying 3 accelerator resources in accelerator set 340 is found to be the container group to be rescheduled. Then in box 418, the operation of the container group to be rescheduled is stopped. For example, stopping the operation of the container group occupying 3 accelerator resources in accelerator set 340 means releasing the accelerator resources occupied by the container group in accelerator set 340 and saving the data state of the container group to shared storage / file system. Subsequently, in box 420, the container group is scheduled to the target node, for example, the container group is scheduled to accelerator set 320. Next, in box 410, resources are allocated. For example, idle resources from accelerator set 340 are allocated to the new container group, and three idle accelerator resources from accelerator set 320 are allocated to the container group awaiting rescheduling, which previously released three accelerator resources from accelerator set 340. Then, in box 416, the container group is restarted from the saved data running state or checkpoint, and the new container group is started simultaneously. This accelerator resource allocation method fully utilizes fragmented accelerator resources in the cluster without waiting for other accelerator resources to be released, thereby maximizing the utilization rate of accelerator resources.

[0060] In some embodiments, a request for allocation of accelerator resources is received in block 404, for example, the new container group requests allocation of 4 accelerator resources. Then, in block 406, it is determined whether there are whole machine accelerator resources in the accelerator set that can satisfy the request. For example, if there are no whole machine accelerator resources in accelerator sets 310, 320, 330, and 340 that can satisfy the request, then in block 408, it is determined whether there are fragmented accelerator resources that can satisfy the request. For example, if there are 6 remaining accelerator resources in accelerator sets 310, 320, 330, and 340 that can satisfy the request, then in block 412, a container group to be scheduled is found. For example, a container group occupying 1 accelerator resource in accelerator set 330 is found as the container group to be scheduled. Next, in box 418, the operation of the container group to be rescheduled is stopped. For example, stopping the operation of the container group occupying one accelerator resource in accelerator set 320 releases the accelerator resource in accelerator set 330 occupied by the container group, and the running data state of the container group is saved to shared storage / file system. Then, in box 420, the container group is scheduled to the target node, for example, the container group is scheduled to accelerator set 330. Subsequently, in box 410, resources are allocated, for example, the idle resources in accelerator set 320 are allocated to the new container group, and one idle accelerator resource in accelerator set 330 is allocated to the container group that previously released one accelerator resource in accelerator set 320. Next, in box 416, the container group is restarted at the saved running data state, and a new container group is started simultaneously to execute the task request.

[0061] In some embodiments, the container group to be scheduled is a container group that executes background tasks or training tasks. Background tasks refer to processes provided by the system that can run in the background, without affecting the execution and provisioning of online services even if the application has been suspended or is no longer running. In some embodiments, the container group to be scheduled has the same resource request priority as the new container group, and this priority is higher than the resource request priority of other container groups besides the container group to be scheduled and the new container group.

[0062] In some embodiments, the nearest deployment node can be selected as the target node for the container to be scheduled, realizing proximity routing capability and thus reducing network overhead during scheduling. In some embodiments, when there are multiple container groups to be scheduled, the container group with the shortest scheduling latency is determined as the container group to be scheduled. For example, if the time to schedule the container group to a node in accelerator set 320 and restore the container group's operation is longer than the time to schedule the container group to a node in accelerator set 330 and restore the container group's operation, then the node in accelerator set 330 is determined as the target node for the container group to be scheduled. In this way, the latency impact of container services can be reduced, improving the user experience.

[0063] Figure 5 A schematic diagram of the architecture 500 of an example migration container group of this disclosure is shown. Figure 5 As shown, both host 570 and host 580 are configured with multiple GPUs and several Remote Direct Access Network Interface Controllers (RNICs) at the hardware level. The RNICs enable Direct Memory Access (DMA) when two or more computers communicate, allowing direct access from one host's memory to another's memory. RNICs also allow data transfer between service and storage configurations. For example, host 580, configured with RNIC 584-1 and RNIC 584-2, can achieve direct memory access with host 570, which is also configured with RNIC 574-1 and RNIC 574-2, avoiding unnecessary latency. In some embodiments, the GPUs can be homogeneous or heterogeneous accelerators.

[0064] refer to Figure 5 Hosts 570 and 580 have host operating system kernels 550 and 560 respectively at the software level. Within kernels 550 and 560, kernel-based Virtual Machine Module (KVM) modules 552 and 562 are configured respectively. KVM is a full virtualization solution based on an operating system kernel and employing hardware-assisted virtualization technology. In the KVM module, virtual machines are implemented as regular operating system processes and scheduled by a standard operating system scheduler. Each virtual CPU (Central Processing Unit) of the virtual machine is implemented as a regular operating system process, allowing KVM to utilize the existing functionality of the operating system kernel.

[0065] Continue to refer to Figure 5 Each host is configured with an Elastic Compute Service (ECS), such as ECS 510, ECS 520, ECS 530, and ECS 540. Each ECS contains one or more container groups, such as container group 512, container group 522, and container group 532. Each container group contains MaaS, such as MaaS 514, MaaS 524, and MaaS 534. An ECS is a resource collection consisting of CPU, memory, cloud disks, etc. MaaS is a model-as-a-service model that integrates the development, deployment, operation, and management of models on a unified platform, such as a container cluster platform, based on the cloud, and provides it to users. This allows users to easily use and manage models without needing to concern themselves with the implementation details and underlying technologies of the models.

[0066] In some embodiments, the ECS can be replaced by a bare metal server (BMS), where the bare metal server occupies the entire host (not shown in the figure). A BMS is a hardware device that combines the characteristics of a traditional physical server with the virtualization service functions of cloud computing technology, and is a product of the combination of hardware and software advantages.

[0067] A container cluster platform is used to manage containerized applications across multiple hosts on a cloud platform. In a container cluster platform, a container group (POD) is the smallest basic unit of deployment and management within the cluster. A container group encapsulates one or more containers, storage resources, an independent network address, and policy options for managing and controlling how containers run. Each container is isolated from the others and has its own file system. A container contains the application and its runtime dependencies, allowing programs to run in a relatively independent environment.

[0068] The container service uses ECS as its underlying resource, enabling the one-click construction of highly available container clusters in the cloud. In some embodiments, a single node is an ECS instance, allowing users to flexibly choose deployment methods based on their business needs. In some embodiments, there are no restrictions on the node type in the container service; nodes can be 32-bit operating systems (x86), heterogeneous, or bare metal. For example, the node configured on host 570 is a bare metal node, while the node configured on host 580 is a cloud server node. In some embodiments, a single container can use one or more accelerator resources on the same node. For example, containers in container group 512 can use GPU 572-1 on host 570. For example, containers in container group 522 can use GPU 572-2 or GPU 572-3 on host 570. For example, containers in container group 532 can use GPU 582-1, GPU 582-2, or GPU 583-1, or any combination thereof, on host 580.

[0069] Continue to refer to Figure 5 Before migrating container group 512 and its Maas 514 instances from ECS 510 to ECS 540, the status files of the data running within the container group need to be saved to shared storage / file system 590. In some embodiments, shared storage / file system 590 can be a distributed file system, including local and remote file systems. In some embodiments, the container group to be scheduled is a container group executing background tasks or training tasks. For example, container group 512 is a container group executing background tasks or training tasks. In this way, scheduling behavior can be avoided from affecting the execution of foreground tasks and disrupting normal program processes. In some embodiments, if container group 512 is a container group executing network-hosted services or in a no-service state, it can be directly shut down.

[0070] In some embodiments, a sidecar container is configured within the container group. A sidecar is a design pattern that decouples application functionality from the application itself as a separate process, allowing for non-intrusive addition of features to the application without adding extra code to meet third-party requirements. In the software architecture, the sidecar is attached to the main application, or parent application, to extend / enhance functionality, while maintaining loose coupling between the sidecar and the main application. If the sidecar container provides an executable waiting for it to become ready, that executable can be invoked in a post-start hook of the container to prevent the startup of other containers in the container group. Through the sidecar approach, the scheduler can interact with daemons within the container group.

[0071] In some embodiments, before scheduling container group 512 from ECS 510 to ECS 540, the scheduler sends a command to stop the operation of container group 512 to the daemon process within container group 512. Upon receiving the command, the daemon process within container group 512 stops the operation of container group 512 and saves the running status of relevant data to shared storage / file system 590. Through this interactive operation, on the one hand, it effectively avoids the degraded user experience of the container group to be rescheduled starting from scratch, and on the other hand, it can make full use of the fragmented accelerator resources in the cluster.

[0072] Figure 6 A schematic diagram illustrating the interaction 600 between the scheduler and the daemon within the container group, representing some embodiments of this disclosure, is shown. For example... Figure 6 As shown, cloud server 610 is configured with container groups 612, 614, and 616. Each container group contains a daemon process, which runs continuously in the background and automatically exits when the main thread ends. The daemon processes can interact with the scheduler to control the container groups. For example, scheduler 620 can interact with daemon process 617 in container group 616, daemon process 615 in container group 614, and daemon process 613 in container group 612. Scheduler 620 sends a stop command to notify the container group to stop running. Upon receiving the command, the daemon process within the container group instructs the daemon process to save the container group's running status data to the shared file system and then stops the container group's operation, before sending a notification message to the scheduler to inform it that the above operations have been completed. The scheduler is the component responsible for application scheduling, scheduling containers to run on worker nodes by configuring node or container group affinity, etc.

[0073] In some embodiments, the scheduler 620 sends a stop command to notify the container group 616 to stop running. Upon receiving the command, the daemon process 617 within the container group 616 saves the running status data of the container group 616 to the shared file system, stops the running of the container group 616, and then sends a notification message to the scheduler 620 to inform it that the above operations have been completed. In some embodiments, the scheduler 620 monitors the status information of all accelerator resources in the cluster, i.e., whether the accelerator resources are idle or occupied, and can update the status information of the accelerator resources after they are released or occupied.

[0074] Figure 7 A block diagram of an apparatus 700 for scheduling accelerator resources, according to some embodiments of the present disclosure, is shown. Figure 7 As shown, the apparatus 700 includes a request acquisition module 702, configured to acquire an accelerator resource allocation request related to a machine learning model, the accelerator resource allocation request indicating the amount of accelerator resources required by a container group running the machine learning model. The apparatus 700 also includes a container group determination module 704, configured to determine a container group to be rescheduled according to a rescheduling policy in response to a rescheduling condition triggered by the accelerator resource allocation request. The apparatus 700 also includes a container group migration module 706, which allocates accelerator resources on the first node for the accelerator resource allocation request by migrating the container group to be rescheduled from a first node to a second node.

[0075] Figure 8 Block diagrams of electronic devices 800 according to some embodiments of the present disclosure are shown. Device 800 may be the device or apparatus described in the embodiments of the present disclosure. Figure 8 As shown, device 800 includes a central processing unit (CPU) and / or a graphics processing unit (GPU) 801, which can perform various appropriate actions and processes according to computer program instructions stored in read-only memory (ROM) 802 or loaded from storage unit 816 into random access memory (RAM) 803. The CPU / GPU 801, ROM 802, and RAM 803 are interconnected via bus 804. Input / output (I / O) interface 805 is also connected to bus 804. Although not shown in... Figure 8 As shown, device 800 may also include a coprocessor.

[0076] Multiple components in device 800 are connected to I / O interface 805, including: input unit 806, such as keyboard, mouse, etc.; output unit 807, such as various types of monitors, speakers, etc.; storage unit 808, such as disk, optical disk, etc.; and communication unit 809, such as network card, modem, wireless transceiver, etc. Communication unit 809 allows device 800 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0077] The various methods or processes described above can be executed by CPU / GPU 801. For example, in some embodiments, the methods can be implemented as computer software programs tangibly contained in a machine-readable medium, such as storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed on device 800 via ROM 802 and / or communication unit 809. When the computer program is loaded into RAM 803 and executed by CPU / GPU 801, one or more steps or actions in the methods or processes described above can be performed.

[0078] In some embodiments, the methods and processes described above can be implemented as a computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for performing various aspects of this disclosure.

[0079] Computer-readable storage media can be tangible devices capable of holding and storing instructions for use by an instruction execution device. Computer-readable storage media can be, for example, but not limited to, electrical storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any suitable combination thereof. The computer-readable storage media used herein are not to be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.

[0080] The computer-readable program instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, a local area network (LAN), a wide area network (WAN), and / or a wireless network, to an external computer or external storage device. The network may include copper cables, fiber optic cables, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage media in the respective computing / processing device.

[0081] Computer program instructions used to perform the operations of this disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, status setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​and conventional procedural programming languages. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry, such as programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), is personalized by utilizing the status information of the computer-readable program instructions to implement various aspects of this disclosure.

[0082] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0083] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0084] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of devices, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0085] The various embodiments of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical applications, or technical improvements to the technology in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.

[0086] The following are some example implementations of this disclosure.

[0087] Example 1. A method for scheduling accelerator resources, comprising:

[0088] Obtain an accelerator resource allocation request related to a machine learning model, the accelerator resource allocation request indicating the amount of accelerator resources required for the container group running the machine learning model;

[0089] In response to the accelerator resource allocation request triggering a rescheduling condition, the container group to be rescheduled is determined according to the rescheduling policy; and

[0090] Accelerator resources are allocated on the first node for the accelerator resource allocation request by migrating the container group to be rescheduled from the first node to the second node.

[0091] Example 2. According to the method of Example 1, the rescheduling condition includes at least one of the following:

[0092] The number of accelerator resources required to run the container group of the machine learning model is greater than the remaining accelerator resources on each node; and

[0093] The number of accelerator resources required to run the container group of the machine learning model is less than or equal to the total remaining amount of accelerator resources on each node.

[0094] Example 3. The method according to any one of Examples 1-2, wherein determining the group of containers to be scheduled according to the rescheduling policy includes:

[0095] The container group to be rescheduled and the first node are determined in response to the fact that the amount of accelerator resources available on one or more nodes is greater than or equal to the amount of accelerator resources occupied by the container group to be rescheduled, and in response to the fact that the amount of accelerator resources available on one or more nodes after the amount of accelerator resources occupied by the container group to be rescheduled is released is greater than or equal to the amount of accelerator resources required by the container group running the machine learning model.

[0096] Example 4. According to the method of any one of Examples 1-3, determine a plurality of candidate container groups, said plurality of container groups including a first container group and a second container group; and

[0097] In response to determining that the restart duration of the second container group is greater than the restart duration of the first container group, the first container group is determined to be the container group to be rescheduled.

[0098] Example 5. The method according to any one of Examples 1-4, wherein the container group to be rescheduled is not the container group executing the front-end task.

[0099] Example 6. The method according to any one of Examples 1-5 further includes:

[0100] In response to the fact that the number of accelerator resources required to run the machine learning model in the container group is greater than the remaining accelerator resources on the one or more nodes, one or more new nodes are configured.

[0101] Example 7. The method according to any one of Examples 1-6, wherein migrating the group of containers to be rescheduled from the first node to the second node comprises:

[0102] Obtain a stop instruction for the container group to be rescheduled, the stop instruction instructing the container group to be rescheduled to stop running;

[0103] In response to receiving the stop command, the running status data of the container group to be rescheduled is saved; and

[0104] In response to receiving a notification indicating that the runtime status data of the container group to be rescheduled has been saved to the file system, the operation of the container group to be rescheduled is stopped.

[0105] Example 8. The method according to any one of Examples 1-7 further includes:

[0106] Start the container group running the machine learning model on the first node; and

[0107] On the second node, allocate accelerator resources to the container group to be rescheduled.

[0108] Example 9. The method according to any one of Examples 1-8, wherein accelerator resources are allocated to the container group to be rescheduled on the second node based on the priority of the accelerator resource allocation request for the container group to be rescheduled being higher than the priority of other container groups on the second node.

[0109] Example 10. The method according to any one of Examples 1-9 further includes:

[0110] In response to a container group requiring less than or equal to the remaining accelerator resources on one or more nodes, accelerator resources are allocated to the accelerator resource allocation request on one of the one or more nodes.

[0111] Example 11. An apparatus for scheduling accelerator resources, comprising:

[0112] The request acquisition module is configured to acquire accelerator resource allocation requests related to a machine learning model, the accelerator resource allocation requests indicating the amount of accelerator resources required for a container group to run the machine learning model;

[0113] A container group determination module is configured to determine the container group to be rescheduled according to a rescheduling policy in response to a rescheduling condition triggered by the accelerator resource allocation request; and

[0114] The container group migration module is configured to allocate accelerator resources on the first node for the accelerator resource allocation request by migrating the container group to be rescheduled from the first node to the second node.

[0115] Example 12. The apparatus according to Example 11, wherein the rescheduling condition includes at least one of the following:

[0116] The number of accelerator resources required to run the container group of the machine learning model is greater than the remaining accelerator resources on each node; and

[0117] The number of accelerator resources required to run the container group of the machine learning model is less than or equal to the total remaining amount of accelerator resources on each node.

[0118] Example 13. The apparatus according to any one of Examples 11-12, wherein the container group determining module comprises:

[0119] The first node determination module is configured to determine the container group to be rescheduled and the first node in response to a situation where the amount of accelerator resources available on one or more nodes is greater than or equal to the amount of accelerator resources occupied by the container group to be rescheduled, and in response to a situation where the amount of accelerator resources available on the one or more nodes after the amount of accelerator resources occupied by the container group to be rescheduled is released is greater than or equal to the amount of accelerator resources required by the container group running the machine learning model.

[0120] Example 14. The apparatus according to any one of Examples 11-13 further includes:

[0121] A candidate container group determination module is configured to determine a plurality of candidate container groups, the plurality of container groups including a first container group and a second container group; and

[0122] The container group to be rescheduled module is configured to determine the first container group as the container group to be rescheduled in response to determining that the restart duration of the second container group is greater than the restart duration of the first container group.

[0123] Example 15. The apparatus according to any one of Examples 11-14, wherein the container group to be rescheduled is not a container group performing a front-end task.

[0124] Example 16. The apparatus according to any one of Examples 11-15 further includes:

[0125] The node configuration module is configured to configure one or more new nodes in response to a situation where the number of accelerator resources required to run the machine learning model in the container group exceeds the remaining amount of accelerator resources on the one or more nodes.

[0126] Example 17. The apparatus according to any one of Examples 11-16, wherein the container group migration module comprises:

[0127] The instruction acquisition module is configured to acquire a stop instruction for the container group to be rescheduled, the stop instruction instructing the container group to be rescheduled to stop running;

[0128] A data storage module is configured to, in response to receiving the stop command, save the running status data of the container group to be rescheduled; and

[0129] The stop module is configured to stop the operation of the container group to be rescheduled in response to receiving a notification indicating that the running status data of the container group to be rescheduled has been saved to the file system.

[0130] Example 18. The apparatus according to Examples 11-17 further includes:

[0131] The startup module is configured to launch and run the container group of the machine learning model on the first node; and

[0132] The resource allocation module is configured to allocate accelerator resources to the container group to be rescheduled on the second node.

[0133] Example 19. An apparatus according to any one of Examples 11-18, wherein accelerator resources are allocated to the container group to be rescheduled on the second node based on the priority of the accelerator resource allocation request for the container group to be rescheduled being higher than the priority of other container groups on the second node.

[0134] Example 20. The apparatus according to any one of Examples 11-19 further includes:

[0135] The resource allocation module is configured to allocate accelerator resources in one of the nodes in response to a container group running the machine learning model having less than or equal to the remaining accelerator resources on one or more nodes.

[0136] Example 21. An electronic device comprising:

[0137] Processor; and

[0138] A memory coupled to the processor, the memory having instructions stored therein, which, when executed by the processor, cause the electronic device to perform actions, the actions including:

[0139] Obtain an accelerator resource allocation request related to a machine learning model, the accelerator resource allocation request indicating the amount of accelerator resources required for the container group running the machine learning model;

[0140] In response to the accelerator resource allocation request triggering a rescheduling condition, the container group to be rescheduled is determined according to the rescheduling policy; and

[0141] Accelerator resources are allocated on the first node for the accelerator resource allocation request by migrating the container group to be rescheduled from the first node to the second node.

[0142] Example 22. The electronic device according to Example 21, wherein the rescheduling condition includes at least one of the following:

[0143] The number of accelerator resources required to run the container group of the machine learning model is greater than the remaining accelerator resources on each node, and

[0144] The number of accelerator resources required to run the container group of the machine learning model is less than or equal to the total remaining accelerator resources on each node.

[0145] Example 23. An electronic device according to any one of Examples 21-22, wherein determining the container group to be scheduled according to a rescheduling policy includes:

[0146] The container group to be rescheduled and the first node are determined in response to the fact that the amount of accelerator resources available on one or more nodes is greater than or equal to the amount of accelerator resources occupied by the container group to be rescheduled, and in response to the fact that the amount of accelerator resources available on the one or more nodes after the amount of accelerator resources occupied by the container group to be rescheduled is released is greater than or equal to the amount of accelerator resources required by the container group running the machine learning model.

[0147] Example 24. The electronic device according to any one of Examples 21-23, wherein the operation further includes:

[0148] Determine multiple candidate container groups, including a first container group and a second container group; and

[0149] In response to determining that the restart duration of the second container group is greater than the restart duration of the first container group, the first container group is determined to be the container group to be rescheduled.

[0150] Example 25. An electronic device according to any one of Examples 21-24, wherein the container group to be rescheduled is not a container group performing a front-end task.

[0151] Example 26. The electronic device according to any one of Examples 21-25, wherein the operation further includes:

[0152] In response to the fact that the number of accelerator resources required to run the machine learning model in the container group is greater than the remaining accelerator resources on the one or more nodes, one or more new nodes are configured.

[0153] Example 27. An electronic device according to any one of Examples 21-26, wherein migrating the group of containers to be rescheduled from the first node to the second node comprises:

[0154] Obtain a stop instruction for the container group to be rescheduled, the stop instruction instructing the container group to be rescheduled to stop running;

[0155] In response to receiving the stop command, the running status data of the container group to be rescheduled is saved; and

[0156] In response to receiving a notification indicating that the runtime status data of the container group to be rescheduled has been saved to the file system, the operation of the container group to be rescheduled is stopped.

[0157] Example 28. An electronic device according to any one of Examples 21-27, wherein the action further includes:

[0158] Start the container group running the machine learning model on the first node; and

[0159] On the second node, allocate accelerator resources to the container group to be rescheduled.

[0160] Example 29. An electronic device according to any one of Examples 21-28, wherein accelerator resources are allocated to the container group to be rescheduled on the second node based on the priority of the accelerator resource allocation request for the container group to be rescheduled being higher than the priority of other container groups on the second node.

[0161] Example 30. An electronic device according to any one of Examples 21-29, wherein the action further includes:

[0162] In response to a container group requiring less than or equal to the remaining accelerator resources on one or more nodes, accelerator resources are allocated to the accelerator resource allocation request on one of the one or more nodes.

[0163] Example 31. A computer-readable storage medium having stored thereon computer-executable instructions, wherein the computer-executable instructions are executed by a processor to implement the method according to any one of Examples 1 to 10.

[0164] Example 32. A computer program product tangibly stored on a computer-readable medium and comprising computer-executable instructions that, when executed by a device, cause the device to perform the method according to any one of Examples 1 to 10.

[0165] Although this disclosure has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.

Claims

1. A method for scheduling accelerator resources, comprising: Obtain an accelerator resource allocation request related to a machine learning model. The accelerator resource allocation request indicates the number of accelerator resources required by a container group to run the machine learning model. The machine learning task corresponding to the accelerator resource allocation request is allocated to the same node. Different nodes are configured with different accelerator clusters. The container group uses one or more accelerator resources of the same node. The accelerator resources belong to the accelerator cluster configured at the hardware level. In response to the accelerator resource allocation request triggering a rescheduling condition, a container group to be rescheduled is determined according to a rescheduling policy. The rescheduling condition is triggered based on the relationship between the number of accelerator resources required by the accelerator resource allocation request and the number of accelerator resources in the accelerator cluster. The rescheduling policy is used to determine the container group to be rescheduled based on this relationship. Accelerator resources are allocated on the first node for the accelerator resource allocation request by migrating the container group to be rescheduled from the first node to the second node.

2. The method of claim 1, wherein the rescheduling condition includes at least one of the following: The number of accelerator resources required to run the container group of the machine learning model is greater than the remaining accelerator resources on each node; and The number of accelerator resources required to run the container group of the machine learning model is less than or equal to the total remaining amount of accelerator resources on each node.

3. The method according to claim 1, wherein determining the container group to be scheduled according to the rescheduling policy includes: The container group to be rescheduled and the first node are determined in response to the fact that the amount of accelerator resources available on one or more nodes is greater than or equal to the amount of accelerator resources occupied by the container group to be rescheduled, and in response to the fact that the amount of accelerator resources available on one or more nodes after the amount of accelerator resources occupied by the container group to be rescheduled is released is greater than or equal to the amount of accelerator resources required by the container group running the machine learning model.

4. The method according to claim 3, further comprising: Determine multiple candidate container groups, including a first container group and a second container group; as well as In response to determining that the restart duration of the second container group is greater than the restart duration of the first container group, the first container group is identified as the container group to be rescheduled.

5. The method of claim 3, wherein the container group to be rescheduled is not the container group executing the front-end task.

6. The method according to claim 3, further comprising: In response to the fact that the number of accelerator resources required to run the machine learning model in the container group is greater than the remaining accelerator resources on the one or more nodes, one or more new nodes are configured.

7. The method of claim 1, wherein migrating the group of containers to be rescheduled from the first node to the second node comprises: Obtain a stop instruction for the container group to be rescheduled, the stop instruction instructing the container group to be rescheduled to stop running; In response to receiving the stop command, the running status data of the container group to be rescheduled is saved; as well as In response to receiving a notification indicating that the runtime status data of the container group to be rescheduled has been saved to the file system, the operation of the container group to be rescheduled is stopped.

8. The method according to claim 1, further comprising: Start the container group that runs the machine learning model on the first node; as well as On the second node, allocate accelerator resources to the container group to be rescheduled.

9. The method of claim 8, wherein the allocation of accelerator resources to the container group to be rescheduled on the second node is based on the priority of the accelerator resource allocation request for the container group to be rescheduled being higher than the priority of other container groups on the second node.

10. The method according to claim 1, further comprising: In response to a container group requiring less than or equal to the remaining accelerator resources on one or more nodes, accelerator resources are allocated to the accelerator resource allocation request on one of the one or more nodes.

11. An apparatus for scheduling accelerator resources, comprising: The request acquisition module is configured to acquire accelerator resource allocation requests related to machine learning models. The accelerator resource allocation requests indicate the number of accelerator resources required by the container group running the machine learning model. The machine learning tasks corresponding to the accelerator resource allocation requests are allocated to the same node. Different nodes are configured with different accelerator clusters. The container group uses one or more accelerator resources of the same node. The accelerator resources belong to the accelerator cluster configured at the hardware level. A container group determination module is configured to, in response to the accelerator resource allocation request triggering a rescheduling condition, determine the container group to be rescheduled according to a rescheduling policy. The rescheduling condition is triggered based on the relationship between the number of accelerator resources required by the accelerator resource allocation request and the number of accelerator resources in the accelerator cluster. The rescheduling policy is used to determine the container group to be rescheduled based on the relationship between the number of accelerator resources required by the accelerator resource allocation request and the number of accelerator resources in the accelerator cluster. The container group migration module is configured to allocate accelerator resources on the first node for the accelerator resource allocation request by migrating the container group to be rescheduled from the first node to the second node.

12. An electronic device, comprising: processor; as well as A memory coupled to the processor, the memory having instructions stored therein, which, when executed by the processor, cause the electronic device to perform the method according to any one of claims 1 to 10.

13. A computer-readable storage medium having stored thereon computer-executable instructions, wherein the computer-executable instructions are executed by a processor to implement the method according to any one of claims 1 to 10.