A deployment method for implementing distributed training using heterogeneous GPUs on k8s

By defining the distributed training task manifest file CR1 and creating several CR2 and Pods on k8s, the problem that the existing technology cannot use multiple heterogeneous GPUs at the same time is solved, and the effect of fully utilizing heterogeneous GPU resources in the same training task is achieved.

CN119718554BActive Publication Date: 2025-06-24KYLIN CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510227978.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-06-24
Estimated Expiration
2045-02-28

AI Technical Summary

Technical Problem

The prior art cannot simultaneously use multiple heterogeneous GPU resources in the k8s cluster in the same training task, resulting in the inability to fully utilize heterogeneous GPUs to accelerate distributed training.

Method used

By defining the distributed training task manifest file CR1 on k8s and creating several CR2 and Pods according to its configuration, the combination of parameter distribution and Pod templates for different GPUs is implemented, allowing multiple heterogeneous GPUs to be used simultaneously in a training task.

Benefits of technology

The purpose of using different GPUs in the same training task is achieved, making full use of heterogeneous GPU resources in the k8s cluster, and accelerating the completion of distributed training tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119718554B_ABST
    Figure CN119718554B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of computers, and provides a deployment method for implementing distributed training using heterogeneous GPUs on k8s, including: writing a distributed training task manifest file CR1 and submitting it to k8s; a first controller monitors the creation of CR1, creates several CR2s according to the configuration of CR1, and assigns training task parameters to each CR2; a second controller monitors the creation of each CR2, creates several Pods according to the configuration of each CR2, and assigns training task parameters to each Pod; the second controller determines whether the current distributed training is elastic training, and if so, creates an Hpa resource for each CR2; after the current distributed training is completed, the first controller deletes all CR2s of the current distributed training task to delete all Pods. It solves the technical problem that it is impossible to use multiple GPUs simultaneously in a training task to accelerate distributed training in the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and specifically provides a deployment method for implementing distributed training using heterogeneous GPUs on k8s. Background Art

[0002] With the rapid development of cloud-native technologies represented by kubernetes, more and more manufacturers choose to run AI training tasks on kubernetes. And due to the increasing size of models, a single GPU can no longer load the entire model, and distributed training has become an inevitable trend. Therefore, how to quickly submit a distributed AI training task on k8s in a simpler way has become a major problem.

[0003] In the existing solutions, distributed AI training task deployment tools represented by kubeflow's training-operator solve this problem. The training-operator continuously monitors the creation of CR resources it is concerned about, creates Pods using the Pod templates defined in the CR, and configures different environment variables for the created Pods according to the parameters required by the training framework (such as Pytorch), indicating that they belong to a specific distributed training task, so that these Pods can jointly complete the training.

[0004] However, from the dominance of Nvidia GPUs to the large-scale deployment of various types of GPUs in actual production tasks, the AI computing power in more data centers maintains a phenomenon of coexistence of multiple GPUs. How to make full use of these heterogeneous GPU resources to accelerate the training of a single large model has become a new problem. For example, a certain k8s cluster has m A-class GPUs and n B-class GPUs. When running a distributed training task on this k8s cluster, the current technology can at most use m A-class GPUs or n B-class GPUs simultaneously, and there is no deployment system that can quickly run a training task to use all m + n GPUs at the same time.

[0005] To use multiple heterogeneous GPUs in the same training task, one cannot use only one Pod template. Because one Pod template can only be configured with one type of GPU and can only be configured with one container image. And in order to adapt to different GPUs, the container images are often different. In addition, for Pod templates using different GPUs, the number of GPUs in a Pod generated by this template may also be different, so the resources such as CPU and memory required by a Pod may also be different. In the existing solutions, kubeflow's training-operator only supports configuring one Pod template and cannot configure multiple GPUs simultaneously, so it cannot use multiple GPUs simultaneously.

[0006] The existing training-operator of Kubeflow realizes the simple and fast submission of a training task in Kubernetes (k8s), but it only supports configuring one Pod template for one training task and cannot make full use of heterogeneous GPU resources in the k8s cluster to accelerate training in the same training task.

[0007] Correspondingly, a new deployment solution is needed in this field to solve the above problems. Summary of the Invention

[0008] In order to overcome the above defects, the present invention is proposed to solve the technical problem that the prior art cannot realize the use of multiple GPUs simultaneously in one training task to accelerate distributed training.

[0009] The present invention provides a deployment method for realizing distributed training using heterogeneous GPUs on k8s, including the steps of:

[0010] Writing a distributed training task manifest file CR1 according to the requirements of the training framework and submitting it to k8s;

[0011] Controller 1 listens for the creation of CR1, creates several CR2s according to the configuration of CR1, and assigns training task parameters to each CR2;

[0012] Controller 2 listens for the creation of each CR2, creates several Pods according to the configuration of each CR2, and assigns training task parameters to each Pod;

[0013] Controller 2 determines whether the current distributed training is elastic training. If so, creates a Horizontal Pod Autoscaler (Hpa) resource for each CR2;

[0014] After the current distributed training is completed, Controller 1 deletes all CR2s of the current distributed training task to delete all Pods.

[0015] Furthermore, it further includes the steps of:

[0016] Controller 1 creates a headless service for the current distributed training, which is used for communication between Pods through domain names.

[0017] Furthermore, it further includes the steps of:

[0018] Controller 1 uses the k8s informer to monitor the updates of each CR2 and updates the status of each CR2 to CR1.Status. Controller 2 uses the k8s informer to monitor the updates of Pods and updates the status of each Pod to CR2.Status, where CR1.Status represents the status of CR1 and CR2.Status represents the status of CR2.

[0019] Furthermore, CR1 includes a Pod.Spec field one and a list field. The Pod.Spec field one is used to configure the fields supported by any k8s Pod template, and the list field is used to configure the GPU names, number of nodes, and images required for different types of GPUs in the training framework requirements.

[0020] Furthermore, the controller one listens for the creation of CR1, and creates several CR2s according to the configuration of CR1 and assigns training task parameters to each CR2, including the steps of:

[0021] The controller one uses the k8s informer to listen for the creation of CR1;

[0022] Traverse the images, GPU names, and number of nodes of different types of GPUs in the CR1 list field, and combine them with the Pod.Spec field one of CR1 to form the Pod.Spec field two of CR2 corresponding to different types of GPUs,

[0023] The controller one calculates the local range of NodeRank for each GPU according to the total number of nodes and the number of nodes of each GPU, determines the CPU and memory of each GPU, and updates them to the Pod.Spec field two of each CR2;

[0024] Create each CR2 according to the generated Pod.Spec field two.

[0025] Furthermore, the controller two listens for the creation of each CR2, and creates several Pods according to the configuration of each CR2 and assigns training task parameters to each Pod, including the steps of:

[0026] The controller two uses the k8s informer to listen for the creation of each CR2;

[0027] The controller two creates different Pod templates, and the number of Pod templates is the same as the number of CR2s;

[0028] The controller two creates the Pod.Spec field three corresponding to the Pod of each CR2 based on the Pod template;

[0029] Create each Pod according to the generated Pod.Spec field three, and use the SetControllerReference method to set the Owner of each Pod to the corresponding CR2.

[0030] Further, the second controller creates the Pod.Spec field three corresponding to each Pod of each CR2 based on the Pod template, including: for each CR2, the second controller creates the Pod.Spec field three of each Pod based on the Pod template, assigns a NodeRank value to each Pod, determines the CPU and memory of each Pod, and updates them into the Pod.Spec field three of each Pod.

[0031] Further, the second controller determines whether the current distributed training is elastic training, including determining whether the number of nodes in CR2 is an interval value by the second controller. If it is an interval value, the current distributed training is elastic training.

[0032] The working principle and beneficial effects of the present invention:

[0033] In implementing the technical solution of the present invention, the first controller creates a number of CR2s by parsing the instance CR1, and the second controller creates a number of Pods by parsing the CR2s. In a distributed training task, the parameters required for distributed training are assigned to each Pod by means of parameter distribution, and the Pods created using different templates are combined into the same training task, achieving the purpose of simultaneously using different GPUs in the same training task. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Referring to the accompanying drawings, the disclosure of the present invention will become more readily understood. It is easily understood by those skilled in the art that these drawings are only for illustrative purposes and are not intended to limit the protection scope of the present invention. In addition, similar numbers in the figures are used to represent similar components, where:

[0035] Figure 1 is a schematic flow chart of a deployment method for implementing distributed training using heterogeneous GPUs on k8s according to the present invention;

[0036] Figure 2 is a schematic flow chart of an implementation manner of the distributed training deployment method according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0037] Some embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood by those skilled in the art that these embodiments are only used to explain the technical principle of the present invention and are not intended to limit the protection scope of the present invention.

[0038] Figure 1 is a schematic flow chart of a deployment method for implementing distributed training using heterogeneous GPUs on k8s according to the present invention, Figure 2 is a schematic flow chart of an implementation manner of the distributed training deployment method according to the present invention. As Figures 1-2As shown, a deployment method for implementing distributed training using heterogeneous GPUs on k8s in this embodiment mainly includes the following steps S1 - S5.

[0039] S1: Write a distributed training task manifest file CR1 according to the requirements of the training framework, and submit CR1 to k8s using kubectl or other tools.

[0040] In one implementation, according to the requirements of the training framework, such as the need for all types of GPUs (graphics processing units), the number of nodes and images required for each type of GPU, etc., write a distributed training task manifest file CR1 (i.e., custom resource). CR1 includes a Pod.Spec field (i.e., container group template, hereinafter referred to as Pod.Spec field one) and a list field. The Pod.Spec field one is used to configure the fields that any k8s Pod template supports configuration; the list field is used to configure the different fields required for different types of GPUs in the training framework requirements, such as GPU name, number of nodes, image, etc.

[0041] In the present invention, the list field and the Pod.Spec field one adopt the way of defining fields, and there is no need to configure a complete Pod template for each type of GPU, which simplifies the complexity of configuring the training task manifest.

[0042] S2: Controller one listens for the creation of CR1, and creates several CR2s according to the configuration of CR1 and assigns training task parameters to each CR2.

[0043] In one implementation, the specific steps of step S2 include the following S21 - S23.

[0044] S21: Controller one uses k8s informer to listen for the creation of CR1.

[0045] S22: Traverse the fields (image, GPU name, number of nodes) corresponding to different types of GPUs in the CR1 list field, and combine them with the Pod.Spec field one of CR1 to form the Pod.Spec field (i.e., container group template, hereinafter referred to as Pod.Spec field two) of the CR2 (i.e., custom resource) corresponding to different types of GPUs.

[0046] This step corresponds to Figure 1 and Figure 2 in " Figure 1 and Figure 2 The "Pod.Spec" in the text boxes pointed to by the digital arrow 4 and the digital arrow 5 in

[0047] Obtain the images, GPU names, and number of nodes of each GPU through the list field. At the same time, Controller 1 calculates the local range of NodeRank for each GPU based on the total number of nodes and the number of nodes of each GPU, determines the CPU and memory of each GPU, and updates them to the second field of Pod.Spec of each CR2.

[0048] S23: Create each CR2 according to the generated second field of Pod.Spec.

[0049] S3: Controller 2 listens for the creation of each CR2, creates several Pods according to the configuration of each CR2, and assigns training task parameters to each Pod.

[0050] In one implementation, the specific steps of step S3 include the following S31 - S33.

[0051] S31: Controller 2 uses k8s informer to listen for the creation of each CR2;

[0052] S32: Controller 2 creates different Pod templates, and the number of Pod templates is the same as the number of CR2s;

[0053] S33: Controller 2 creates the Pods of each CR2 based on the Pod templates;

[0054] Specifically, for each CR2, Controller 2 creates the Pod.Spec field of each Pod (i.e., the container group template, hereinafter referred to as the third field of Pod.Spec) based on the Pod template, assigns a NodeRank value to each Pod, determines the CPU and memory of each Pod, and updates them to the third field of Pod.Spec of each Pod;

[0055] Figure 2 In, "parameter1 - 1", "parameter1 - 2", "parameter2 - 1", "parameter2 - 2" correspond to the parameter configurations in the third field of Pod.Spec, that is, the NodeRank value, CPU, and memory of each Pod, etc. The second field of Pod.Spec stores the local range of NodeRank (such as 0 - 1), while the third field of Pod.Spec stores the specific values of 0 and 1.

[0056] S34: Create each Pod (i.e., container group) according to the generated third field of Pod.Spec, and use the SetControllerReference method to set the Owner of each Pod to the corresponding CR2.

[0057] Among them, the number of Pods created by each CR2 is the same as the number of nodes of the GPU corresponding to each CR2.

[0058] Furthermore, Controller 1 creates a headless service for the current distributed training, which is used for communication between Pods through domain names. In k8s, the IP address of a Pod is unstable and cannot be obtained before the Pod starts, while the domain name is stable and can be prefetched.

[0059] In this embodiment, Controller 1 uses k8s informer to monitor the updates of each CR2, and updates the status of each CR2 (the status update of CR2 is completed by Controller 2) to CR1.Status. Controller 2 uses k8s informer to monitor the updates of the Pod status, and updates the status of each Pod to CR2.Status. The update of the Pod status is completed by k8s. CR1.Status represents the status of CR1, and CR2.Status represents the status of CR2.

[0060] S4: Controller 2 determines whether the current distributed training is elastic training. If so, it creates an Hpa resource for each CR2.

[0061] In one implementation, Controller 2 determines whether the number of nodes of CR2 is an interval value. If it is an interval value, that is, the current distributed training is elastic training, it creates an Hpa resource for this CR2 so that it can automatically scale in and out.

[0062] The number of nodes of CR2 is the same as the number of nodes of the GPU. However, the number of nodes of CR2 is separate from the second field of Pod.Spec. The number of nodes is not saved in the second field of Pod.Spec, but in a field of CR2.

[0063] S5: After the current distributed training is completed (successfully or failed), Controller 1 deletes all CR2s of the current distributed training task to delete all Pods.

[0064] In this embodiment, when creating a Pod, the SetControllerReference method is used to set the Owner of the Pod to CR2. Utilizing the cascading deletion ability of k8s, when CR2 is deleted, the Pods belonging to it will also be automatically deleted, preventing completed or failed tasks from continuing to occupy resources.

[0065] Based on the above steps S1 - S5, Controller 1 creates several CR2s by parsing instance CR1, and Controller 2 creates several Pods by parsing CR2s. In a distributed training task, different Pod templates are constructed for different GPUs through parameter distribution, and Pods corresponding to the number of nodes of the corresponding GPUs under the Pod templates are created through different Pod templates. Parameters required for distributed training are allocated to each Pod, and the Pods created using different templates are combined into the same training task, achieving the purpose of simultaneously using different GPUs in the same training task. Then, the overall state is reflected through local states to obtain the final result of the training task, and finally, resources are recycled after the task is completed, forming a closed loop.

[0066] In addition, in the present invention, Controller 1 coordinates the parameter distribution of the entire training task, and Controller 2 further performs parameter distribution and creates training task Pods. The two - level controller mode can make the overall system more decoupled and is convenient for creating Hpa for elastic training.

[0067] It should be noted that although the above - mentioned embodiments describe each step in a specific order, those skilled in the art can understand that for the purpose of implementing the present invention, different steps do not necessarily have to be executed in such an order. They can be executed simultaneously (in parallel) or in other orders, and these variations are all within the protection scope of the present invention.

[0068] Here, some terms related to the present invention are explained.

[0069] K8s: Also known as Kubernetes, abbreviated as k8s, is an open - source container orchestration and management tool.

[0070] Pod: The smallest control unit of k8s, used to carry and run actual application programs. When running a distributed training task in k8s, one Pod is a node of a training task.

[0071] Pod.Spec: The Pod template, which can be used to create Pods.

[0072] AI: Artificial Intelligence.

[0073] Nvidia: A GPU manufacturer.

[0074] GPU: Graphics Processing Unit, mainly used for graphics rendering and parallel computing, and mainly used for parallel computing in the field of AI.

[0075] Kubeflow: A set of tools for the development, training, optimization, deployment, and management of machine learning on the k8s platform.

[0076] Training-operator: A component in Kubeflow used to quickly deploy distributed training tasks on Kubernetes (k8s).

[0077] CRD: Custom Resource Definition, used to extend k8s resources when native k8s resources are insufficient to meet specific requirements.

[0078] CR: Custom Resource, an instance of CRD. Multiple CRs can be created for one CRD.

[0079] Operator: An application written to solve a k8s-oriented problem, usually consisting of one or more controllers and one or more CRDs.

[0080] CR1 and CR2: Respectively represent an instance of CRD1 corresponding to Controller 1 and an instance of CRD2 corresponding to Controller 2. Among them, CR1 is created by the user, and CR2 is created by Controller 1.

[0081] CR1.Status and CR2.Status: Respectively represent the status of CR1 and the status of CR2, which are part of CR1 and CR2. They exist when CR1 and CR2 are created, do not need to be created separately, are not assigned values by the user, are controlled by the controller, and users do not need to pay attention when creating.

[0082] Informer: A core component in the k8s client library, used to monitor changes in k8s API resources and notify users when resources change.

[0083] Headless Service: Headless server, used to provide a stable access entry for Pods.

[0084] Kubectl: A k8s client command-line tool.

[0085] Hpa: A component in k8s used to elastically scale in and out Pods based on resources such as CPU and memory.

[0086] Pytorch: An open-source deep learning framework known for its excellent flexibility and ease of use.

[0087] NodeRank: A parameter assigned to each node in Pytorch distributed training, which is a value from 0 to the number of nodes minus 1 in k8s (i.e., for each Pod).

[0088] So far, the technical solution of the present invention has been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it is easy for those skilled in the art to understand that the protection scope of the present invention is obviously not limited to these specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will fall within the protection scope of the present invention.

Claims

1. A deployment method for implementing distributed training using heterogeneous GPUs on k8s, characterized in that: Includes steps: Write the distributed training task list file CR1 according to the training framework requirements and submit it to k8s; Controller 1 monitors the creation of CR1, creates several CR2s according to the configuration of CR1, and assigns training task parameters to each CR2; Controller 2 monitors the creation of each CR2, creates several Pods according to the configuration of each CR2, and assigns training task parameters to each Pod; Controller 2 determines whether the current distributed training is elastic training. If so, it creates an Hpa resource for each CR2. After the current distributed training is completed, the controller deletes all CR2s of the current distributed training task to delete all Pods; CR1 includes a Pod.Spec field 1 and a list field. The Pod.Spec field 1 is used to configure the fields supported by any k8s Pod template, and the list field is used to configure the GPU name, number of nodes, and image required by different types of GPUs in the training framework requirements. Controller 1 monitors the creation of CR1, creates several CR2s according to the configuration of CR1, and assigns training task parameters to each CR2, including the following steps: Controller 1 uses k8s informer to monitor the creation of CR1; Traverse the images, GPU names, and node numbers of different types of GPUs in the CR1 list field, and combine them with the Pod.Spec field of CR1 to form the Pod.Spec field 2 of CR2 corresponding to different types of GPUs. Controller 1 calculates the local range of NodeRank of each GPU based on the total number of nodes and the number of nodes of each GPU, determines the CPU and memory of each GPU, and updates them to the Pod.Spec field 2 of each CR2; Create each CR2 according to the generated Pod.Spec field 2; CR1, CR2: represent the instance of CRD1 corresponding to controller 1 and the instance of CRD2 corresponding to controller 2, respectively. CR1 is created by the user, and CR2 is created by controller 1.

2. According to claim 1, a deployment method for implementing distributed training using heterogeneous GPUs on k8s is characterized in that: Also includes the steps: Controller 1 creates a headless service for the current distributed training, which is used for communication between Pods through domain names.

3. According to claim 1, a deployment method for implementing distributed training using heterogeneous GPUs on k8s is characterized in that: Also includes the steps: Controller 1 uses k8s informer to monitor updates of each CR2 and updates the status of each CR2 to CR1.Status. Controller 2 uses k8s informer to monitor updates of Pod and updates the status of each Pod to CR2.Status. CR1.Status represents the status of CR1, and CR2.Status represents the status of CR2.

4. According to claim 1, a deployment method for implementing distributed training using heterogeneous GPUs on k8s is characterized in that: Controller 2 monitors the creation of each CR2, creates several Pods according to the configuration of each CR2, and assigns training task parameters to each Pod, including the following steps: Controller 2 uses k8s informer to monitor the creation of each CR2; Controller 2 creates different Pod templates. The number of Pod templates is the same as that of CR2. Controller 2 creates Pod.Spec field 3 corresponding to each CR2 Pod based on the Pod template; Create each Pod according to the generated Pod.Spec field 3, and use the SetControllerReference method to set the Owner of each Pod to the corresponding CR2.

5. According to claim 4, a deployment method for implementing distributed training using heterogeneous GPUs on k8s is characterized in that: The controller 2 creates the Pod.Spec field 3 corresponding to each Pod of CR2 based on the Pod template, including: for each CR2, the controller 2 creates the Pod.Spec field 3 of each Pod based on the Pod template, assigns a NodeRank value to each Pod, determines the CPU and memory of each Pod, and updates them to the Pod.Spec field 3 of each Pod.

6. According to claim 1, a deployment method for implementing distributed training using heterogeneous GPUs on k8s is characterized in that: The controller 2 determines whether the current distributed training is elastic training, including the controller 2 determining whether the number of nodes in CR2 is an interval value. If it is an interval value, the current distributed training is elastic training.

Citation Information

Patent Citations

  • Elastic distributed training method in deep learning scene

    CN114756385A

  • Multi-cluster model training method and device, equipment and medium

    CN116450355A

  • Heterogeneous parallel computing system and distributed training method

    CN118796402A