GPU cluster efficient scheduling method and system oriented to algorithm training
By adopting an efficient scheduling method based on kubernetes in GPU clusters, the problems of low resource utilization and low task execution efficiency in the existing technology are solved, and efficient and stable execution of algorithm training is achieved, which is suitable for the rapid iteration needs of the autonomous driving industry.
Patent Information
- Application Number
- CN202510056861.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-05-13
AI Technical Summary
The existing technology cannot effectively schedule and allocate machine learning algorithm tasks in GPU clusters, resulting in low resource utilization and low task execution efficiency. Especially in the autonomous driving industry, it is necessary to quickly iterate and efficiently train algorithm models.
A GPU cluster efficient scheduling method for algorithm training is proposed. Based on kubernetes and its ecological chain, it realizes accurate scheduling and efficient execution of tasks through training task splitting, cluster utilization monitoring, video memory sharing and isolation, and task scheduling.
Through task splitting and precise scheduling, the efficiency and resource utilization of algorithm training are improved, the task is ensured to be executed quickly and stably, the risk of cluster instability is reduced, and the overall performance of the GPU cluster is improved.
Smart Images

Figure CN119987967A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology and task scheduling technology, and in particular to a GPU cluster efficient scheduling method and system for algorithm training. Background Art
[0002] With the rise of the autonomous driving industry chain and the prevalence of technologies such as artificial intelligence and machine learning, the current autonomous driving industry often encounters situations where a large number of machine learning algorithm tasks need to be run on GPU server clusters. Since these tasks themselves are extremely time-consuming and consume computing resources, how to quickly obtain results to strengthen the autonomous driving model has become a core issue of concern to the industry. On the one hand, we can take the approach of increasing costs and increasing the size of the cluster; in addition, how to improve the overall utilization efficiency of a fixed number of clusters and the utilization rate of each machine is also a crucial breakthrough point. Since the container management system kubernetes was made public, there has been a new way to automatically schedule and allocate machine learning tasks on a large scale. In the future, the industry is likely to use this system in large quantities to enable large-scale server cluster scheduling and operation and maintenance, thereby bringing flexibility and efficiency improvements to the research and development and iteration of autonomous driving technology. In the existing technology, Kubernetes itself has a completely exclusive scheduling strategy for GPUs. Once a GPU is occupied by a task, it cannot be shared by other tasks. The industry's k8s cluster based on docker as the container-runtime has realized GPU memory sharing, which can enable multiple tasks to share the same GPU graphics card on the same node. For example, there is a node with two GPU cards (GPU1 and GPU2, both with 15GiB video memory), and at a certain moment two tasks are using GPU1, at this time it can be considered that the two tasks are sharing the GPU. For example, CN111506404A discloses a shared GPU scheduling method based on Kubernetes, and its technical solution is to propose this solution based on the kubernetes framework, which implements solutions such as GPU sharing and video memory isolation to improve GPU utilization in the industry. The following common strategies are used: 1. Binpack strategy, which will give priority to scheduling Pods to graphics card nodes with fewer remaining GPU resources when there are multiple graphics cards, and schedule to new graphics cards when the remaining resources of this graphics card are insufficient, to ensure that the graphics card resources are compactly occupied, to avoid fragmentation, and to ensure that resources are available when there are new large GPU resource requests; 2. Spread strategy, which will evenly distribute Pods to multiple graphics cards when there are multiple graphics cards, and if the number of Pods is the same as the number of graphics cards, each Pod should be assigned to an independent graphics card to make full use of graphics card resources and avoid waste. 3. Exclusive strategy: This strategy will only select graphics cards that have not been allocated when scheduling Pods, ensuring that the Pod can use this graphics card alone to avoid interference from other Pods. It is used in situations where graphics card resources are highly occupied. This technical route ensures that the risk of task failure is reduced, and improves GPU utilization and isolation through GPU sharing and video memory isolation. However, the above strategy only discusses how GPUs should be allocated to graphics cards from a macro perspective, and does not have any effect on GPU sharing.The essence of the patented method of sharing GPUs is to allocate GPUs according to the above strategy based on the resources claimed by the tasks, without any restrictions or supervision on the behavior after allocation. Without an effective scheduling allocation strategy, and without sorting resource requirements according to different types of tasks, it is impossible to ensure that tasks are scheduled to the most suitable nodes. In addition, video memory isolation refers to the isolation and restriction of the graphics card memory used by each task when multiple tasks share the same GPU. For example. Figure 3 As shown in the figure, the left part is the existing technology. The video memory requested by the Pod is not restricted, and all Pods share the 15G video memory of the GPU. Figure 3 As shown in the left figure, the video memory sizes requested by the tasks executed by pod1 and pod2 are 2Gi and 3Gi respectively, but in fact, the existing memory that each task can use is the entire video memory capacity of GPU1 (15Gi). This is the case where there is no video memory isolation. If the task is not set up with video memory isolation, it will face the risk of running failure. For example, a task requests 2Gi of video memory, but actually uses 15Gi (occupies the entire GPU), then all other tasks that share the GPU will fail to run. The situation of video memory isolation is as follows Figure 3 As shown in the figure on the right, the video memory used by each task is exclusive and limited. For example, if a pod applies for 2Gi video memory, it can only use 2Gi video memory. If the task uses 3Gi video memory, the program will exit abnormally, while other tasks sharing this GPU will not be affected. However, the solution of the above patent only records the allocation information as a fixed value at the time of allocation. The system is completely unaware of whether the actual operation exceeds these allocated resource limits, and it is impossible to make any response.
[0003] With the release of Kubernetes and the support of GPU scheduling by major frameworks such as Pytorch, TensorFlow, and NVIDIA, large-scale GPU clusters dedicated to machine learning algorithm model training have emerged. The autonomous driving industry is the most typical scenario. To truly implement autonomous driving, it is not only necessary to collect a large amount of data, but most importantly, it is necessary to clean and train these large amounts of raw data through algorithm models to obtain useful data so that the recognition capabilities of artificial intelligence can be quickly iterated. Among them, how to make the algorithm tasks accurately scheduled and executed as quickly, efficiently and stably as possible has become the most critical part of the entire autonomous driving implementation process. Summary of the invention
[0004] In order to solve the above problems, this proposal proposes an efficient GPU cluster scheduling method for algorithm training, which can accurately schedule algorithm tasks and complete them as quickly, efficiently and stably as possible. In order to achieve the above objectives, the present invention adopts the following technical solutions:
[0005] An efficient GPU cluster scheduling method for algorithm training based on Kubernetes and its ecosystem includes the following steps:
[0006] S1, training task splitting, using data parallelism to split the data into multiple parts and execute them on different GPUs. The model executed on each GPU is exactly the same;
[0007] S2, cluster utilization and health indicator monitoring, obtains GPU and cluster indicator information from Prometheus, combines scripts to monitor the status of GPU cluster servers in real time, captures these indicators and stores them in the database, and calculates the hardware parameters and overall utilization of the cluster in real time based on the indicators of each machine, and then pushes these indicators to the BI system for display;
[0008] S3, video memory sharing and isolation, and GPU time-sharing. After tasks can be scheduled across machines, GPU resources are refined in granularity based on shared GPU video memory and video memory isolation plug-ins, and any part of the video memory of any single GPU is provided to subtasks split by any task.
[0009] S4, task scheduling, records the demand information of submitted tasks according to the demand type of tasks, and sorts the cluster nodes according to different resource calculation methods for tasks with different demands through the cluster node resource information entered in S2. Then, the tasks are scheduled to the nodes with the highest ranking, and the redis distributed lock is used to atomically record the GPU allocation and usage information, and the resource information of the cluster nodes is updated at the same time.
[0010] The present invention describes an intelligent scheduling logic for this scenario, which supports various tasks that need to be clustered, can monitor the cluster status in real time and dynamically divide the task into multiple subtasks according to the requirements and type of the submitted task, and accurately allocate each subtask to the appropriate node to ensure that all subtasks are completed quickly and stably. It can not only achieve an effect similar to the task being completed stably in a complete manner, but also greatly improve the time required for the task execution through task segmentation, while the required resources remain almost unchanged; thereby improving the efficiency of key links such as algorithm training or similar resource-consuming tasks.
[0011] The present invention implements a scheduling strategy for large-scale machine learning clusters, dynamically allocates tasks with the help of the k8s cluster management framework and enters data such as task information, cluster status information, and task real-time status information, combined with open source GPU video memory sharing and isolation plug-ins, inserts custom scheduling logic with the help of k8swebhook, and can achieve each step of each subtask to be scheduled to the appropriate cluster node on demand with the help of plug-ins and custom webhooks and podOperator (controller). The maximum flexibility of algorithm task scheduling and high GPU cluster utilization are achieved, so that users do not need to care about the logic and details of the underlying allocation, but only need to wait for the task to be completed and get the final execution result. There is no situation where the task exceeds the limit and is terminated by the system, and the risk of cluster instability is also reduced.
[0012] Furthermore, the S1 comprises the following steps:
[0013] During the training task splitting process, the results are continuously merged, and finally the result is exactly the same as executing the task on the same GPU.
[0014] Furthermore, the main steps of S1 include (taking 8 GPUs as an example): the current data packet (called the global packet) will be divided into different sub-data packets (called local packets) with the same number as the number of GPUs. Each GPU uses the same model to independently process a sub-data packet: each GPU runs the model forward once, then runs the model backward once, outputs the gradient parameter results, and efficiently aggregates the results obtained by all GPUs. Aggregation is performed after each step of the model is completed to keep all GPUs in a synchronized state.
[0015] Furthermore, in S2, the gpu-index of the container is checked through Docker and aligned with the gpu-index information monitored from Prometheus, and the GPU information occupied by each container is collected, so that containers or tasks that exceed the limit can be removed in real time in subsequent steps.
[0016] Furthermore, in S3, the GPU driver presents a physical GPU as multiple GPU units to the Kubelet on the node, and the Kubelet allocates these GPU units to various containers on demand.
[0017] Furthermore, in S3, if a physical GPU is allocated to multiple containers, the GPU hardware and NVIDIA driver are used to switch contexts between containers based on usage to achieve time-sharing access to the GPU.
[0018] Furthermore, in S3, for NVIDIA Tesla A100 series and later graphics cards from NVIDIA, the officially supported MIG (multi-instance GPU) graphics card function can be used to physically achieve true video memory isolation, that is, the original A100 graphics card can be physically split into multiple sub-graphics cards, which are exclusively used by different tasks or containers.
[0019] Furthermore, in S4, for tasks with high CPU demand, the ratio of CPU resource demand to other resource demand is 2:1:1:1. With the help of node resource monitoring implemented in S2, the current resources of all nodes are calculated according to the formula: CPU idle percentage × 2 + GPU video memory idle percentage × 1 + memory idle percentage × 1 + GPU idle utilization percentage × 1 (here), and then all node resources are ranked according to the result value. Then, during scheduling, the task is split and scheduled to the node with the highest ranking.
[0020] The weight in the formula is determined according to the proportion of resource demand. The resource sorting and scheduling process is completed in the k8s webhook called before each task is assigned, and the scheduler is the pod controller that comes with k8s.
[0021] The present invention also provides a GPU cluster efficient scheduling system for algorithm training, based on Kubernetes and its ecological chain, including:
[0022] The training task splitting module is used to use data parallelism to split the data into multiple parts and assign them to different GPUs for execution. The model executed on each GPU is exactly the same.
[0023] The cluster utilization and health indicator monitoring module is used to obtain GPU and cluster indicator information from Prometheus, and use scripts to monitor the status of GPU cluster servers in real time, capture these indicators and store them in the database. At the same time, the hardware parameters and overall utilization of the cluster are calculated in real time based on the indicators of each machine, and then these indicators are pushed to the BI system for display;
[0024] The memory sharing and isolation and GPU time-sharing sharing modules are used to implement GPU resource granularity refinement based on the shared GPU memory and memory isolation plug-ins after tasks can be scheduled across machines, and to provide any part of the memory of any single GPU to the subtasks split by any task;
[0025] The task scheduling module is used to record the demand information of submitted tasks according to the types of task requirements, and the cluster node resource information entered by the cluster utilization and health indicator monitoring module. For tasks with different types of requirements, the cluster nodes are sorted according to different resource calculation methods, and then the tasks are scheduled to the highest-ranked nodes. The redis distributed lock is used to atomically record the GPU allocation and usage information, and the resource information of the cluster nodes is updated at the same time.
[0026] Furthermore, the training task splitting module includes performing the following tasks:
[0027] During the training task splitting process, the results are continuously merged, and finally the result is exactly the same as executing the task on the same GPU.
[0028] Furthermore, the main working steps of the training task splitting module include: the current data packet (called the global packet) will be divided into different sub-data packets (called local packets) with the same number as the number of GPUs, and each GPU uses the same model to independently process a sub-data packet: each GPU runs the model forward once, and then runs the model backward once, outputs the gradient parameters and other results, and efficiently aggregates the results obtained by all GPUs. Similar aggregation is performed after each step of the model is completed, so that all GPUs are in a synchronized state.
[0029] Furthermore, when executing its tasks, the cluster utilization and health indicator monitoring module checks the gpu-index of the container through docker and aligns it with the gpu-index information monitored from Prometheus, and collects the GPU information occupied by each container, so as to eliminate the containers or tasks that exceed the limit in real time in the subsequent steps.
[0030] Furthermore, during the process of the video memory sharing and isolation and GPU time-sharing sharing modules performing their tasks, the GPU driver presents a physical GPU as multiple GPU units to the Kubelet on the node, and the Kubelet allocates these GPU units to each container as needed.
[0031] Furthermore, in the process of the video memory sharing and isolation and GPU time-sharing sharing modules performing their tasks, if a physical GPU is allocated to multiple containers for use, the GPU hardware and NVIDIA driver are used to switch contexts between the containers according to usage to achieve time-sharing access to the GPU.
[0032] Furthermore, in the process of the video memory sharing and isolation and the GPU time-sharing sharing module performing its tasks, for NVIDIA Tesla A100 series and later graphics cards of NVIDIA Corporation, the function of the officially supported MIG (multi-instance GPU) graphics card can be used to physically achieve true video memory isolation, that is, the original A100 graphics card can be physically split into multiple sub-graphics cards, which are exclusively occupied by different tasks or containers.
[0033] Furthermore, in the process of executing tasks by the task scheduling module, for tasks with high CPU demand, the ratio of CPU resource demand to other resource demand is 2:1:1:1. With the help of node resource monitoring implemented by the cluster utilization and health indicator monitoring module, the current resources of all nodes are calculated according to the formula: CPU idle percentage × 2 + GPU video memory idle percentage × 1 + memory idle percentage × 1 + GPU idle utilization percentage × 1, (here), and then all node resources are ranked according to the result value, and then the task is split during scheduling and scheduled to the node with high ranking;
[0034] The weight in the formula is determined according to the proportion of resource demand. The resource sorting and scheduling process is completed in the k8s webhook called before each task is assigned, and the scheduler is the pod controller that comes with k8s. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 This is a schematic diagram of the task scheduling architecture of the present invention;
[0036] Figure 2 It is a schematic diagram of the method steps of the present invention;
[0037] Figure 3 A schematic diagram for comparing video memory without isolation and video memory isolation;
[0038] Figure 4 A table showing the resource usage of each container; DETAILED DESCRIPTION
[0039] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0040] like Figure 1 and Figure 2As shown in the figure, an efficient GPU cluster scheduling method for algorithm training is based on Kubernetes and its ecosystem, including the following steps:
[0041] S1, training task splitting, using data parallelism to split the data into multiple parts and execute them on different GPUs. The model executed on each GPU is exactly the same;
[0042] S2, cluster utilization and health indicator monitoring, obtains GPU and cluster indicator information from Prometheus, combines scripts to monitor the status of GPU cluster servers in real time, captures these indicators and stores them in the database, and calculates the cluster hardware parameters and overall utilization in real time based on the indicators of each machine, and then pushes these indicators to the BI system for display; in this way, the overall operation of the cluster can be easily observed at the front desk, such as Figure 4 shown.
[0043] S3, video memory sharing and isolation, and GPU time-sharing sharing. After tasks can be scheduled across machines, GPU resources can be refined based on shared GPU video memory and video memory isolation plug-ins, and any part of the video memory of any single GPU can be provided to subtasks divided by any task; thus achieving the maximum degree of freedom in task scheduling. For example, you can apply https: / / github.com / AliyunContainerService / gpushare-scheduler-extender, the latest version of Alibaba's open source GPU scheduler.
[0044] S4, task scheduling, records the demand information of submitted tasks according to the demand type of tasks, and sorts the cluster nodes according to different resource calculation methods for tasks with different demands through the cluster node resource information entered in S2. Then, the tasks are scheduled to the nodes with the highest ranking, and the redis distributed lock is used to atomically record the GPU allocation and usage information, and the resource information of the cluster nodes is updated at the same time.
[0045] The present invention describes an intelligent scheduling logic for this scenario, which supports various tasks that need to be clustered, can monitor the cluster status in real time and dynamically divide the task into multiple subtasks according to the requirements and type of the submitted task, and accurately allocate each subtask to the appropriate node to ensure that all subtasks are completed quickly and stably. It can not only achieve an effect similar to the task being completed stably in a complete manner, but also greatly improve the time required for the task execution through task segmentation, while the required resources remain almost unchanged; thereby improving the efficiency of key links such as algorithm training or similar resource-consuming tasks.
[0046] The present invention implements a scheduling strategy for large-scale machine learning clusters, dynamically allocates tasks with the help of the k8s cluster management framework and enters data such as task information, cluster status information, and task real-time status information, combined with open source GPU video memory sharing and isolation plug-ins, inserts custom scheduling logic with the help of k8s webhook, and can achieve each step of each subtask to be scheduled to the appropriate cluster node on demand, and with the help of plug-ins, custom webhooks, and podOperator (controller). The maximum flexibility of algorithm task scheduling and high GPU cluster utilization are achieved, so that users do not need to care about the logic and details of the underlying allocation, but only need to wait for the task to be completed and get the final execution result. There is no situation where the task exceeds the limit and is terminated by the system, and the risk of cluster instability is also reduced.
[0047] In some embodiments, the S1 comprises the following steps:
[0048] During the training task splitting process, the results are continuously merged, and finally the result is exactly the same as executing the task on the same GPU.
[0049] This method improves GPU utilization efficiency and uses data parallelism to improve task execution efficiency. For a model training task, the data is split into multiple parts and executed on different GPUs, while the model executed on each GPU is exactly the same. The results will be continuously merged during the process, and the final result is exactly the same as the result of executing the task on the same GPU, thus achieving more efficient GPU calls.
[0050] In the specific implementation method, the main steps of S1 include: the current data packet (called the global packet) will be divided into different sub-data packets (called local packets) with the same number as the number of GPUs, and each GPU uses the same model to independently process a sub-data packet: each GPU runs the model forward once, then runs the model backward once, outputs the gradient parameter results, and efficiently aggregates the results obtained by all GPUs. Aggregation is performed after each step of the model is completed, so that all GPUs are in a synchronized state. Take 8GPUs as an example: the current data packet (called the global packet) will be divided into 8 different sub-data packets (called local packets). For example, if the global packet has 512 sampling results, then each local packet will have 64 sampling results. Each GPU uses the same model to independently process a local packet: they will run the model forward once and then run the model backward once, output the gradient parameters and other results. The results obtained by each of the 8 GPUs are efficiently aggregated. Similar aggregation is performed after each step of the model is completed, so all GPUs are always in a synchronized state.
[0051] In some embodiments, in S2, the gpu-index of the container is checked through docker and aligned with the gpu-index information monitored from Prometheus, and the GPU information occupied by each container is collected, so that the containers or tasks that use the excess quota can be removed in real time in the subsequent steps. This ensures that each task can be carried out smoothly and quickly.
[0052] In some embodiments, in S3, the GPU driver presents a physical GPU as multiple GPU units to the Kubelet on the node, and the Kubelet allocates these GPU units to various containers as needed.
[0053] If a physical GPU is allocated to multiple containers, the GPU hardware and NVIDIA driver are used to switch contexts between containers based on usage to achieve time-sharing access to the GPU. The content in the Pod definition file does not need to be changed because kubelet will configure the time-sharing GPU to expose multiple available GPUs.
[0054] In S3, for NVIDIA Tesla A100 series and later graphics cards from NVIDIA, the officially supported MIG (multi-instance GPU) graphics card function can be used to physically achieve true video memory isolation, that is, the original A100 graphics card can be physically split into multiple sub-graphics cards, which are exclusively used by different tasks or containers.
[0055] In some embodiments, in S4, for high CPU demand tasks, the ratio of CPU resource demand to other resource demand is 2:1:1:1. With the help of node resource monitoring implemented in S2, the current resources of all nodes are calculated according to the formula: CPU idle percentage × 2 + GPU video memory idle percentage × 1 + memory idle percentage × 1 + GPU idle utilization percentage × 1, (here), and then all node resources are ranked according to the result value, and then the task is split during scheduling and scheduled to the node with high ranking;
[0056] The weight in the formula is determined according to the proportion of resource demand. The resource sorting and scheduling process is completed in the k8s webhook called before each task is assigned, and the scheduler is the pod controller that comes with k8s.
[0057] The present invention also includes the step of task resource monitoring, which is to monitor the information of task resource occupation in real time through the data information pulled from Prometheus, compare it with the information of task resource application in the redis queue, and obtain the cluster resource information to which the task is scheduled. If the actual occupation of the task exceeds its application resources for a period of time and the resources of the node where the task is located are not available, the task is stopped and kicked out, and added to the waiting execution queue to ensure the normal operation of other tasks and the promised completion characteristics of each task. However, if the task or pod is the exclusive graphics card (including the exclusive graphics card in the case of MIG), there is no need to kick the task out.
[0058] The present invention also provides a GPU cluster efficient scheduling system for algorithm training, based on Kubernetes and its ecological chain, including:
[0059] The training task splitting module is used to use data parallelism to split the data into multiple parts and assign them to different GPUs for execution. The model executed on each GPU is exactly the same.
[0060] The cluster utilization and health indicator monitoring module is used to obtain GPU and cluster indicator information from Prometheus, and use scripts to monitor the status of GPU cluster servers in real time, capture these indicators and store them in the database. At the same time, the hardware parameters and overall utilization of the cluster are calculated in real time based on the indicators of each machine, and then these indicators are pushed to the BI system for display;
[0061] The memory sharing and isolation and GPU time-sharing sharing modules are used to implement GPU resource granularity refinement based on the shared GPU memory and memory isolation plug-ins after tasks can be scheduled across machines, and to provide any part of the memory of any single GPU to the subtasks split by any task;
[0062] The task scheduling module is used to record the demand information of submitted tasks according to the types of task requirements, and the cluster node resource information entered by the cluster utilization and health indicator monitoring module. For tasks with different types of requirements, the cluster nodes are sorted according to different resource calculation methods, and then the tasks are scheduled to the highest-ranked nodes. The redis distributed lock is used to atomically record the GPU allocation and usage information, and the resource information of the cluster nodes is updated at the same time.
[0063] Here’s how it works:
[0064] First, the training task splitting module performs the following tasks:
[0065] Using data parallelism, the data is split into multiple parts and executed on different GPUs. The model executed on each GPU is exactly the same. During the training task splitting process, the results are continuously merged, and finally the result is exactly the same as executing the task on the same GPU.
[0066] The main working steps of the training task splitting module include: the current data packet (called the global packet) will be divided into different sub-data packets (called local packets) with the same number as the number of GPUs, and each GPU uses the same model to independently process a sub-data packet: each GPU runs the model forward once, then runs the model backward once, outputs the gradient parameters and other results, and efficiently aggregates the results obtained by all GPUs. Similar aggregation is performed after each step of the model is completed, so that all GPUs are in a synchronized state.
[0067] At the same time, the cluster utilization and health indicator monitoring module obtains GPU and cluster indicator information from Prometheus while executing its tasks, and uses scripts to monitor the status of the GPU cluster server in real time, captures these indicators and stores them in the database. At the same time, it calculates the hardware parameters and overall utilization of the cluster in real time based on the indicators of each machine, and then pushes these indicators to the BI system for display; after checking the gpu-index of the container through docker, it is aligned with the gpu-index information monitored from Prometheus, and the GPU information occupied by each container is collected, so that containers or tasks that exceed the limit can be eliminated in real time in subsequent steps.
[0068] During the process of the video memory sharing and isolation and GPU time-sharing sharing modules executing their tasks, after the tasks can be scheduled across machines, the granularity of GPU resources is refined according to the shared GPU video memory and video memory isolation plug-ins, and any part of the video memory of any single GPU is provided to the subtasks divided by any task; the GPU driver presents a physical GPU as multiple GPU units to the Kubelet on the node, and the Kubelet allocates these GPU units to each container as needed.
[0069] In the process of the video memory sharing and isolation and GPU time-sharing sharing modules performing their tasks, if a physical GPU is allocated to multiple containers for use, the GPU hardware and NVIDIA driver are used to switch contexts between the containers according to usage to achieve time-sharing access to the GPU.
[0070] In the process of the video memory sharing and isolation and the GPU time-sharing sharing module performing its tasks, for NVIDIA Tesla A100 series and later graphics cards of NVIDIA Corporation, the function of the officially supported MIG (multi-instance GPU) graphics card can be used to physically achieve true video memory isolation, that is, the original A100 graphics card can be physically split into multiple sub-graphics cards, which are exclusively occupied by different tasks or containers.
[0071] Then, the task scheduling module records the demand information of the submitted tasks according to the task requirements, and the cluster node resource information entered by the cluster utilization and health indicator monitoring module. For tasks with different types of requirements, the cluster nodes are sorted according to different resource calculation methods, and then the tasks are scheduled to the highest-ranked nodes. The redis distributed lock is used to atomically record the GPU allocation and usage information, and the resource information of the cluster nodes is updated at the same time.
[0072] For example, in the process of executing tasks by the task scheduling module, for tasks with high CPU demand, the ratio of CPU resource demand to other resource demand is 2:1:1:1. With the help of the node resource monitoring implemented by the cluster utilization and health indicator monitoring module, the current resources of all nodes are calculated according to the formula: CPU idle percentage × 2 + GPU video memory idle percentage × 1 + memory idle percentage × 1 + GPU idle utilization percentage × 1, and then all node resources are ranked according to the result value, and then the task is split and scheduled to the high-ranking node during scheduling; the weight in the formula is determined according to the ratio of resource demand size; the resource sorting and scheduling process is completed in the k8swebhook called before each task assignment, and the scheduler is the pod controller that comes with k8s. After completing the resource sorting of the cluster, the task is scheduled to the node with the highest ranking and the redis distributed lock is used to atomically record the allocation and use information of the GPU, and the resource information of the cluster node is updated at the same time.
[0073] Finally, you can also set up a task resource monitoring module. By pulling data information from Prometheus, you can monitor the information about task resource occupation in real time, compare it with the information about task resource application in the redis queue, and obtain the cluster resource information to which the task is scheduled. If the actual resource occupation of the task exceeds the resource application for a period of time and there are no free resources on the node where the task is located, the task will be stopped and kicked out, and added to the queue waiting for execution to ensure the normal operation of other tasks and the promised completion characteristics of each task. However, if the task or pod is the exclusive graphics card (including the exclusive graphics card in the MIG case), there is no need to kick the task out.
[0074] The above description is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can make equivalent replacements or changes according to the technical scheme and inventive concept of the present invention within the technical scope disclosed by the present invention, which should be covered by the protection scope of the present invention.
Claims
1. An efficient GPU cluster scheduling method for algorithm training based on kubernetes and its ecosystem, characterized in that: The following steps are involved: S1, training task splitting, using data parallelism to split the data into multiple parts and execute them on different GPUs. The model executed on each GPU is exactly the same; S2, cluster utilization and health indicator monitoring, obtains GPU and cluster indicator information from Prometheus, combines scripts to monitor the status of GPU cluster servers in real time, captures these indicators and stores them in the database, and calculates the hardware parameters and overall utilization of the cluster in real time based on the indicators of each machine, and then pushes these indicators to the BI system for display; S3, video memory sharing and isolation, and GPU time-sharing. After tasks can be scheduled across machines, GPU resources are refined in granularity based on shared GPU video memory and video memory isolation plug-ins, and any part of the video memory of any single GPU is provided to subtasks split by any task. S4, task scheduling, records the demand information of submitted tasks according to the demand type of tasks, and sorts the cluster nodes according to different resource calculation methods for tasks with different demands through the cluster node resource information entered in S2. Then, the tasks are scheduled to the nodes with the highest ranking, and the redis distributed lock is used to atomically record the GPU allocation and usage information, and the resource information of the cluster nodes is updated at the same time.
2. The efficient GPU cluster scheduling method for algorithm training according to claim 1 is characterized in that: The S1 comprises the following steps: During the training task splitting process, the results are continuously merged, and finally the result is exactly the same as executing the task on the same GPU.
3. The efficient GPU cluster scheduling method for algorithm training according to claim 2 is characterized in that: The main steps of S1 include: the current data packet will be divided into different sub-data packets with the same number of GPUs, and each GPU will use the same model to independently process a sub-data packet: each GPU runs the model forward once, then runs the model backward once, outputs the gradient parameter results, and efficiently aggregates the results obtained by all GPUs. Aggregation is performed after each step of the model is completed to keep all GPUs in a synchronized state.
4. The efficient GPU cluster scheduling method for algorithm training according to any one of claims 1 to 3, characterized in that: In S2, the gpu-index of the container is checked through Docker and aligned with the gpu-index information monitored from Prometheus, and the GPU information occupied by each container is collected, so that containers or tasks that exceed the limit can be removed in real time in subsequent steps.
5. The efficient GPU cluster scheduling method for algorithm training according to any one of claims 1 to 3, characterized in that: In S3, the GPU driver presents a physical GPU as multiple GPU units to the Kubelet on the node, and the Kubelet allocates these GPU units to individual containers on demand.
6. The efficient GPU cluster scheduling method for algorithm training according to claim 5 is characterized in that: In S3, if a physical GPU is allocated to multiple containers, the GPU hardware and NVIDIA driver are used to switch contexts between containers based on usage to achieve time-sharing access to the GPU.
7. The efficient GPU cluster scheduling method for algorithm training according to claim 6 is characterized in that: In S3, for NVIDIA Tesla A100 series and later graphics cards from NVIDIA, the officially supported MIG (multi-instance GPU) graphics card function can be used to physically achieve true video memory isolation, that is, the original A100 graphics card can be physically split into multiple sub-graphics cards, which are exclusively used by different tasks or containers.
8. The efficient GPU cluster scheduling method for algorithm training according to any one of claims 1 to 3, characterized in that: It also includes a task resource monitoring step. By pulling data information from Prometheus, it monitors the resource occupation information of tasks in real time, compares it with the resource application information of tasks in the redis queue, and obtains the cluster resource information to which the task is scheduled. If the actual resource occupation of the task exceeds its application resource for a period of time and there are no free resources on the node where the task is located, the task will be stopped and kicked out, and added to the queue waiting for execution to ensure the normal operation of other tasks and the promised completion characteristics of each task.
9. The efficient GPU cluster scheduling method for algorithm training according to any one of claims 1 to 3, characterized in that: In S4, for tasks with high CPU demand, the ratio of CPU resource demand to other resource demand is 2:1:1:
1. With the help of node resource monitoring implemented in S2, the current resources of all nodes are calculated according to the formula: CPU idle percentage × 2 + GPU video memory idle percentage × 1 + memory idle percentage × 1 + GPU idle utilization percentage × 1. Then, all node resources are ranked according to the result value, and the task is split and scheduled to the node with high ranking. The weight in the formula is determined according to the proportion of resource demand. The resource sorting and scheduling process is completed in the k8s webhook called before each task is assigned, and the scheduler is the pod controller that comes with k8s.
10. An efficient GPU cluster scheduling system for algorithm training, based on kubernetes and its ecosystem, characterized by: include: The training task splitting module is used to use data parallelism to split the data into multiple parts and assign them to different GPUs for execution. The model executed on each GPU is exactly the same. The cluster utilization and health indicator monitoring module is used to obtain GPU and cluster indicator information from Prometheus, and use scripts to monitor the status of GPU cluster servers in real time, capture these indicators and store them in the database. At the same time, the hardware parameters and overall utilization of the cluster are calculated in real time based on the indicators of each machine, and then these indicators are pushed to the BI system for display; The memory sharing and isolation and GPU time-sharing sharing modules are used to implement GPU resource granularity refinement based on the shared GPU memory and memory isolation plug-ins after tasks can be scheduled across machines, and to provide any part of the memory of any single GPU to the subtasks split by any task; The task scheduling module is used to record the demand information of submitted tasks according to the types of task requirements, and the cluster node resource information entered by the cluster utilization and health indicator monitoring module. For tasks with different types of requirements, the cluster nodes are sorted according to different resource calculation methods, and then the tasks are scheduled to the highest-ranked nodes. The redis distributed lock is used to atomically record the GPU allocation and usage information, and the resource information of the cluster nodes is updated at the same time.
11. The GPU cluster efficient scheduling system for algorithm training according to claim 10, characterized in that: The training task splitting module includes performing the following tasks: During the training task splitting process, the results are continuously merged, and finally the result is exactly the same as executing the task on the same GPU.
12. The GPU cluster efficient scheduling system for algorithm training according to claim 11, characterized in that: The main working steps of the training task splitting module include: the current data packet (called the global packet) will be divided into different sub-data packets (called local packets) with the same number as the number of GPUs, and each GPU uses the same model to independently process a sub-data packet: each GPU runs the model forward once, then runs the model backward once, outputs the gradient parameters and other results, and efficiently aggregates the results obtained by all GPUs. Similar aggregation is performed after each step of the model is completed, so that all GPUs are in a synchronized state.
13. The GPU cluster efficient scheduling system for algorithm training according to any one of claims 10-12, characterized in that: When executing its tasks, the cluster utilization and health indicator monitoring module checks the gpu-index of the container through docker and aligns it with the gpu-index information monitored from Prometheus, collects the GPU information occupied by each container, and removes the containers or tasks that exceed the limit in real time in subsequent steps.
14. The GPU cluster efficient scheduling system for algorithm training according to any one of claims 10-12, characterized in that: During the process of the video memory sharing and isolation and GPU time-sharing sharing modules performing their tasks, the GPU driver presents a physical GPU as multiple GPU units to the Kubelet on the node, and the Kubelet allocates these GPU units to each container as needed.
15. The GPU cluster efficient scheduling method for algorithm training according to claim 14, characterized in that: In the process of the video memory sharing and isolation and GPU time-sharing sharing modules performing their tasks, if a physical GPU is allocated to multiple containers for use, the GPU hardware and NVIDIA driver are used to switch contexts between the containers according to usage to achieve time-sharing access to the GPU.
16. The GPU cluster efficient scheduling system for algorithm training according to claim 15, characterized in that: In the process of the video memory sharing and isolation and the GPU time-sharing sharing module performing its tasks, for NVIDIA Tesla A100 series and later graphics cards of NVIDIA Corporation, the function of the officially supported MIG (multi-instance GPU) graphics card can be used to physically achieve true video memory isolation, that is, the original A100 graphics card can be physically split into multiple sub-graphics cards, which are exclusively occupied by different tasks or containers.
17. The GPU cluster efficient scheduling system for algorithm training according to any one of claims 10-12, characterized in that: In the process of executing tasks by the task scheduling module, for tasks with high CPU demand, the ratio of CPU resource demand to other resource demand is 2:1:1:
1. With the help of node resource monitoring implemented by the cluster utilization and health indicator monitoring module, the current resources of all nodes are calculated according to the formula: CPU idle percentage × 2 + GPU video memory idle percentage × 1 + memory idle percentage × 1 + GPU idle utilization percentage × 1, (here), and then all node resources are ranked according to the result value, and then the task is split during scheduling and scheduled to the node with high ranking; The weight in the formula is determined according to the proportion of resource demand. The resource sorting and scheduling process is completed in the k8s webhook called before each task is assigned, and the scheduler is the pod controller that comes with k8s.
18. The GPU cluster efficient scheduling system for algorithm training according to any one of claims 10-12, characterized in that: It also includes a task resource monitoring module, which is used to monitor the information of task resource occupation in real time through the data information pulled from Prometheus, compare it with the information of task resource application in the redis queue, and obtain the cluster resource information to which the task is scheduled. If the actual resources occupied by the task exceed its applied resources for a period of time and there are no free resources on the node where the task is located, the task will be stopped and kicked out, and added to the waiting execution queue to ensure the normal operation of other tasks and the promised completion characteristics of each task.
Citation Information
Patent Citations
Shared GPU (Graphics Processing Unit) scheduling method based on Kubernetes
CN111506404A