A GPU sharing method, device, equipment and medium
Through CUDA hijacking and MPS technology, the problems of parallel execution and resource isolation on the GPU are solved, and efficient GPU utilization and secure isolation are achieved. Users can choose the execution mode of the Pod to meet different needs.
Patent Information
- Application Number
- CN202211164060.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-23
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2042-09-23
AI Technical Summary
In the prior art, how to implement multiple Pods scheduling on the same GPU during GPU sharing, and multiple Pods can be executed in parallel on the same GPU to improve GPU utilization and ensure resource isolation and security.
CUDA hijacking and MPS technology is used to create pods by obtaining service creation information, and adding environment variable information to the pods, determine the GPU operation mode and sharing method, filter the target GPU nodes, and use CUDA to perform GPU sharing calculations to realize parallel execution or time-sharing scheduling of multiple pods.
It realizes high GPU utilization and video memory isolation of multiple Pods. Users can choose to execute Pods in parallel or schedule in time, make full use of the GPU computing core, or exclusively or share GPUs.
Smart Images

Figure CN115495215B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technologies, and particularly to a GPU sharing method, apparatus, device, and medium. Background Art
[0002] It has become very common to use GPUs (Graphics Processing Units) to provide computing power in machine learning and deep learning. A typical scenario is in a data center, where, relying on Kubernetes (a portable and scalable open-source platform) as a container orchestration environment, a cloud environment cluster is built to deploy machine learning and deep learning services. The nodes (servers) in the cluster are divided into different types. Nodes equipped with GPUs are called GPU nodes, and other nodes are called CPU nodes. GPU nodes are responsible for specific machine learning and deep learning tasks, while CPU nodes are responsible for cluster management, service scheduling, etc. However, since the resources such as video memory, registers, and threads provided by a single GPU are quite sufficient, usually a single Kubernetes Pod cannot fully utilize the video memory, registers, threads, etc. of a single GPU. Therefore, a technology is needed to achieve scheduling multiple Pods of multiple services onto the same GPU, so as to achieve the purpose of high GPU utilization. In addition, the running modes of multiple Pods (the smallest deployable computing unit that can be created and managed in Kubernetes) on the same GPU can be divided into two types. The first running mode is time-sharing scheduling, that is, only one Pod is performing calculations at the same time. One drawback of this mode is that if a Pod cannot utilize all the computing cores of the GPU, then the GPU will be in a low-load state, resulting in waste of resources. The other running mode is the parallel running mode, that is, multiple Pods perform calculations simultaneously, so that the cores of the GPU can be utilized more fully. In the existing technology, rCUDA performs resource isolation through hijacking calls and supports GPU resource pooling at the same time. Pooling means using GPU resources in the form of remote access. The task uses the local CPU and the GPU of another machine, and the two communicate through the network. Also for this reason, the shared module needs to separate the calls of the CPU and the GPU. The mixed-compiled program will insert some non-open-source CUDA APIs. Therefore, the CUDA (a parallel computing platform and application programming interface (API)) provided by the author needs to be used to compile the CPU and GPU parts of the program separately. If this product is used, users need to recompile, which has a certain impact on users. Mig (Multi-Instance GPU) is a GPU sharing solution officially released by NVIDIA. NVIDIA isolates resources at the underlying hardware level and can completely achieve isolation of computing / communication / configuration / errors. It evenly distributes the computing cores and video memory to GPU instances, supporting a maximum of dividing the computing cores into 7 parts and the video memory into 8 parts. Mig can at most divide the computing cores of a single GPU into 7 parts and the video memory into 8 parts, and cannot perform finer-grained partitioning. Moreover, only relatively new models of GPUs support Mig.MPS (Multi-Process Service) is a GPU sharing solution officially released by NVIDIA. It utilizes the HyperQ capability on the GPU, allowing multiple services to share the same GPU context. It permits the computational and data replication operations of Pods from different services to be executed concurrently on the same GPU to maximize GPU utilization. Since MPS (Master Production Schedule) executes Pods from multiple services within the same GPU context, and once a Pod from one service encounters an execution anomaly, this anomaly will be propagated to Pods of other services, MPS cannot provide secure isolation.
[0003] As can be seen from the above, during the process of GPU sharing, how to schedule multiple Pods onto the same GPU and enable multiple Pods to execute in parallel on the same GPU to improve GPU utilization is an issue to be resolved in this field. Summary of the Invention
[0004] In view of this, the purpose of the present invention is to provide a GPU sharing method, apparatus, device, and medium, which can achieve scheduling multiple Pods onto the same GPU, enabling multiple Pods to execute in parallel on the same GPU, and improving GPU utilization. The specific solutions are as follows:
[0005] In the first aspect, the present application discloses a GPU sharing method, including:
[0006] Obtain service creation information, create a Pod based on the service creation information, and add environment variable information to the Pod;
[0007] Determine the GPU operation mode and GPU sharing mode based on the service creation message;
[0008] Determine the target GPU node according to the GPU sharing mode and the GPU operation mode, and send the environment variable information in the Pod to the target GPU node, so that the target GPU node performs GPU sharing calculations on the environment variable information based on the GPU operation mode.
[0009] Optionally, the obtaining of the service creation information includes:
[0010] Detect whether the local service creation is triggered;
[0011] If the local service creation is triggered, obtain the service creation information.
[0012] Optionally, the creating of the Pod based on the service creation information and adding environment variable information to the Pod includes:
[0013] Create Pods with the same number as the number of replicas in the service creation information;
[0014] Add environment variable information to the Pod based on the annotation information in the service creation information.
[0015] Optionally, determining the GPU operation mode and GPU sharing mode based on the service creation message includes:
[0016] Determine the GPU operation mode based on the annotation information in the service creation message; wherein, the GPU operation mode includes a parallel operation mode and a time-sharing scheduling mode;
[0017] Determine the GPU sharing mode based on the annotation information in the service creation message; wherein, the GPU sharing mode includes service-exclusive GPU and service-shared GPU.
[0018] Optionally, determining the target GPU node according to the GPU sharing mode and the GPU operation mode includes:
[0019] If the GPU operation mode is the parallel operation mode, determine whether the GPU sharing value in the annotation information is the same as the pre-obtained service sharing value. If the GPU sharing value in the annotation information is the same as the pre-obtained service sharing value, determine that the GPU sharing mode is service-exclusive GPU, and then screen out the GPU nodes with the GPU node label value of true and the GPU sharing mode of service-exclusive GPU from all GPU nodes as the target GPU nodes;
[0020] If the GPU operation mode is the time-sharing scheduling mode, determine that the GPU sharing mode is service-shared GPU, and then screen out the GPU nodes with the label value of false and the GPU sharing mode of service-shared GPU from all GPU nodes as the target GPU nodes.
[0021] Optionally, sending the environment variable information in the Pod to the target GPU node so that the target GPU node performs GPU sharing calculation on the environment variable information based on the GPU operation mode includes:
[0022] Send the environment variable information in the Pod to the CUDA in the target GPU node so that the CUDA in the target GPU node performs GPU sharing calculation on the environment variable information based on the GPU operation mode.
[0023] Optionally, the CUDA in the target GPU node performs GPU sharing calculation on the environment variable information based on the GPU operation mode, including:
[0024] If the GPU operating mode is a parallel operating mode, the CUDA performs GPU shared computing on the environment variable information by using a preset computing function and a video memory application and release function.
[0025] If the GPU operating mode is a time-sharing scheduling mode, the CUDA determines time ratio information from the service creation information, and performs GPU shared computing on the time ratio information by using the computing function.
[0026] In a second aspect, the present application discloses a GPU sharing device, including:
[0027] A Pod creation module, configured to obtain service creation information, create a Pod based on the service creation information, and add environment variable information to the Pod.
[0028] A GPU operating mode determination module, configured to determine a GPU operating mode and a GPU sharing mode based on the service creation message.
[0029] A GPU shared computing module, configured to determine a target GPU node according to the GPU sharing mode and the GPU operating mode, and send the environment variable information in the Pod to the target GPU node, so that the target GPU node performs GPU shared computing on the environment variable information based on the GPU operating mode.
[0030] In a third aspect, the present application discloses an electronic device, including:
[0031] A memory, configured to store a computer program;
[0032] A processor, configured to execute the computer program to implement the foregoing GPU sharing method.
[0033] In a fourth aspect, the present application discloses a computer storage medium, configured to store a computer program; wherein, when the computer program is executed by a processor, the steps of the foregoing disclosed GPU sharing method are implemented.
[0034] It can be seen that the present application provides a GPU sharing method, which includes obtaining service creation information, creating a Pod based on the service creation information, and adding environment variable information to the Pod; determining a GPU operation mode and a GPU sharing mode based on the service creation message; determining a target GPU node according to the GPU sharing mode and the GPU operation mode, and sending the environment variable information in the Pod to the target GPU node, so that the target GPU node performs GPU sharing calculation on the environment variable information based on the GPU operation mode. Based on CUDA hijacking and MPS technology, the present application can realize scheduling multiple Pods to the same GPU, so as to achieve the purpose of high GPU utilization rate. Moreover, the video memories of multiple Pods on the same GPU are isolated. Users can choose to execute multiple Pods on the same GPU in parallel to make full use of the computing cores of the GPU, or schedule multiple Pods at different times and execute them alternately. They can also choose to exclusively occupy the GPU for the service, or share the GPU with other services. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained according to the provided drawings.
[0036] Figure 1 It is a flowchart of a GPU sharing method disclosed in the present application;
[0037] Figure 2 It is a flowchart of a GPU sharing method disclosed in the present application;
[0038] Figure 3 It is a specific architecture diagram of a GPU sharing method disclosed in the present application;
[0039] Figure 4 It is a flowchart of CUDA hijacking disclosed in the present application;
[0040] Figure 5 It is a schematic structural diagram of a GPU sharing device disclosed in the present application;
[0041] Figure 6 It is a structural diagram of an electronic device provided by the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0042] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0043] It has become very common to use GPUs (Graphics Processing Units) to provide computing power in machine learning and deep learning. A typical scenario is in a data center, where, relying on Kubernetes (a portable and scalable open-source platform) as a container orchestration environment, a cloud environment cluster is built to deploy machine learning and deep learning services. The nodes (servers) in the cluster are divided into different types. Nodes equipped with GPUs are called GPU nodes, and other nodes are called CPU nodes. GPU nodes are responsible for specific machine learning and deep learning tasks, while CPU nodes are responsible for cluster management, service scheduling, etc. However, since the video memory, registers, threads, etc. provided by a single GPU are very sufficient, usually a single Kubernetes Pod cannot fully utilize the video memory, registers, threads, etc. of a single GPU. Therefore, a technology is needed to achieve scheduling multiple Pods of multiple services onto the same GPU, so as to achieve the purpose of high GPU utilization. In addition, the running modes of multiple Pods (the smallest deployable computing unit that can be created and managed in Kubernetes) on the same GPU can be divided into two types. The first running mode is the time-sharing scheduling mode, that is, only one Pod is performing calculations at the same time. One drawback of this mode is that if a Pod cannot utilize all the computing cores of the GPU, then the GPU will be in a low-load state, resulting in waste of resources. The other running mode is the parallel running mode, that is, multiple Pods perform calculations simultaneously, so that the cores of the GPU can be utilized more fully. In the prior art, rCUDA uses hijacking calls for resource isolation and also supports GPU resource pooling. Pooling means using GPU resources in the form of remote access. The task uses the local CPU and the GPU of another machine, and the two communicate through the network. Also for this reason, the shared module needs to separate the calls of the CPU and the GPU. The mixed-compiled program will insert some non-open-source CUDA APIs. Therefore, the CUDA (a parallel computing platform and application programming interface (API)) provided by the author needs to be used to compile the CPU and GPU parts of the program separately. If this product is used, users need to recompile, which has a certain impact on users. Mig (Multi-Instance GPU) is a GPU sharing solution officially released by NVIDIA. NVIDIA isolates resources at the underlying hardware level and can completely achieve isolation of computing / communication / configuration / errors. It evenly distributes computing cores and video memory to GPU instances, supporting a maximum of dividing computing cores into 7 parts and video memory into 8 parts. Mig can at most divide the computing cores of a single GPU into 7 parts and the video memory into 8 parts, and cannot perform finer-grained partitioning. Moreover, only relatively new models of GPUs support Mig.MPS (Multi-Process Service) is a GPU sharing solution officially released by NVIDIA. It utilizes the Hyper Q capability on the GPU, allowing multiple services to share the same GPU context, enabling the concurrent execution of computing and data replication operations of Pods from different services on the same GPU to maximize GPU utilization. Since MPS (Master Production Schedule) executes Pods from multiple services in the same GPU context, and once a Pod from one service encounters an execution exception, this exception will be propagated to Pods from other services, MPS cannot provide secure isolation. As can be seen from the above, in the process of GPU sharing, how to schedule multiple Pods to the same GPU and enable multiple Pods to execute in parallel on the same GPU to improve GPU utilization is an issue to be solved in this field. Based on CUDA hijacking and MPS technology, this application can schedule multiple Pods to the same GPU, thereby achieving the goal of high GPU utilization. Moreover, the video memory of multiple Pods on the same GPU is isolated. Users can choose to execute multiple Pods from the same service in parallel on the same GPU, fully utilizing the computing cores of the GPU, or schedule multiple Pods time-divisionally and execute them alternately. Users can also choose for a service to exclusively occupy the GPU or share the GPU with other services.
[0044] See Figure 1 As shown, an embodiment of the present invention discloses a GPU sharing method, which may specifically include:
[0045] Step S11: Obtain service creation information, create a Pod based on the service creation information, and add environment variable information to the Pod.
[0046] In this embodiment, it is detected whether local service creation is triggered. If local service creation is triggered, service creation information is obtained, and then a Pod is created based on the service creation information, and environment variable information is added to the Pod.
[0047] Step S12: Determine the GPU operation mode and GPU sharing mode based on the service creation message.
[0048] Step S13: Determine the target GPU node according to the GPU sharing mode and the GPU operation mode, and send the environment variable information in the Pod to the target GPU node, so that the target GPU node performs GPU sharing calculation on the environment variable information according to the GPU operation mode.
[0049] In this embodiment, after determining the target GPU node according to the GPU sharing mode and the GPU operation mode, the environment variable information in the Pod is sent to the CUDA in the target GPU node, so that the CUDA in the target GPU node performs GPU sharing calculation on the environment variable information based on the GPU operation mode.
[0050] Specifically, if the GPU operation mode is the parallel operation mode, the CUDA performs GPU sharing calculation on the environment variable information by using a preset calculation function and a video memory application and release function. If the GPU operation mode is the time-sharing scheduling mode, the CUDA determines the time ratio information from the service creation information and performs GPU sharing calculation on the time ratio information by using the calculation function.
[0051] In this embodiment, service creation information is obtained, a Pod is created based on the service creation information, and environment variable information is added to the Pod; the GPU operation mode and the GPU sharing mode are determined based on the service creation message; a target GPU node is determined according to the GPU sharing mode and the GPU operation mode, and the environment variable information in the Pod is sent to the target GPU node, so that the target GPU node performs GPU sharing calculation on the environment variable information based on the GPU operation mode. Based on CUDA hijacking and MPS technology, this application can schedule multiple Pods to the same GPU, thereby achieving the purpose of high GPU utilization. Moreover, the video memories of multiple Pods on the same GPU are isolated. Users can choose to execute multiple Pods on the same GPU in parallel, making full use of the computing cores of the GPU, or multiple Pods can be scheduled in a time-sharing manner and executed alternately. They can also choose for the service to exclusively occupy the GPU or share the GPU with other services.
[0052] See Figure 2 As shown, an embodiment of the present invention discloses a GPU sharing method, which may specifically include:
[0053] Step S21: Obtain service creation information, create Pods with the same number as the number of replicas in the service creation information based on the number of replicas in the service creation information, and add environment variable information to the Pods based on the annotation information in the service creation information.
[0054] In this embodiment, when it is detected that a service is created in the cluster, the corresponding number of Pods and some other network-related resources are created according to the replica count in the service yaml file. When creating a Pod, an annotation (gpusharing / mps) in the service yaml is used to determine whether to inject the environment variable "MPS" into the container of the Pod. The "MPS" environment variable determines whether to schedule the Pod of this service to a certain GPU on a GPU node in parallel running mode or to a certain GPU on a GPU node in time-sharing scheduling mode. Among them, the GPU operation modes include parallel running mode and time-sharing scheduling mode.
[0055] Step S22: Determine the GPU operation mode based on the annotation information in the service creation message; among them, the GPU operation modes include parallel running mode and time-sharing scheduling mode, and determine the GPU sharing mode based on the annotation information in the service creation message; among them, the GPU sharing modes include exclusive GPU for the service and shared GPU for the service.
[0056] Specifically, if the GPU operation mode is parallel running mode, determine whether the GPU sharing value in the annotation information is the same as the service sharing value obtained in advance. If the GPU sharing value in the annotation information is the same as the service sharing value obtained in advance, determine that the GPU sharing mode is exclusive GPU for the service, and then screen out the GPU nodes with the GPU node label value of true and the GPU sharing mode of exclusive GPU for the service from all GPU nodes as the target GPU nodes; if the GPU operation mode is time-sharing scheduling mode, determine that the GPU sharing mode is shared GPU for the service, and then screen out the GPU nodes with the label value of false and the GPU sharing mode of shared GPU for the service from all GPU nodes as the target GPU nodes.
[0057] Step S23: Determine the target GPU node according to the GPU sharing mode and the GPU operation mode, and send the environment variable information in the Pod to the target GPU node, so that the target GPU node performs GPU sharing calculation on the environment variable information based on the GPU operation mode.
[0058] The GPU sharing system (GPUSharing) architecture of this application is as Figure 3As shown, the master node: the management node of the Kubernetes cluster, which includes a controller module and a scheduler module; the controller: creates corresponding Pods and other resources for the services created by the user, and injects the following environment variables into the containers of the Pods: MPS, which identifies whether the service uses MPS; the scheduler: schedules the Pod corresponding to the service to a certain GPU and creates the corresponding vGPU; the GPU node: the computing node in the Kubernetes cluster with a GPU installed. There is a node agent running on such a node, and the node agent further includes the following sub-modules: the configuration client: responsible for registering the GPU resources of this node with the scheduler, and at the same time writing the list information of the Pods running on the GPU into a file; the hijacking scheduler: reads the list information of the Pods written by the configuration client and completes the following functions: in the time-sharing scheduling mode, it is responsible for allocating time slices to the Pods and restricting the video memory allocation; in the parallel running mode, it is responsible for restricting the video memory allocation. When the scheduler schedules a newly created Pod, it will consider the following two conditions: according to the annotation (gpusharing / mps) of this Pod to decide whether to schedule this Pod to a certain GPU on the GPU node in the parallel running mode or to a certain GPU on the time-sharing scheduling GPU node; according to the annotation (gpusharing / group) of this Pod to decide whether the service of this Pod exclusively occupies the GPU or shares the same GPU with other services. If this annotation is not set in the yaml file of the service, then this service will share the same GPU with other services; if this annotation is set, then only the services with the same value of this annotation will share the same GPU; from the above explanation, it can be seen that if we want a service to exclusively occupy a GPU, then we only need to make its "gpusharing / group" annotation unique (the meaning of exclusive occupation is: for example, if service 1 has two Pods, denoted as Pod 1 and Pod 2, and it is known that Pod 1 is scheduled to GPU node 1, then only Pods of service 1 can be on GPU node 1, and Pods of other services cannot be scheduled to GPU node 1. That is to say, in the future, only Pod 2 may be scheduled to GPU node 1, and the same goes for GPU node 2). Only the GPU nodes that meet both of the above two conditions will become the target GPU nodes. There are two types of GPU nodes in the cluster. The first type is the node with MPS enabled. This type of node will have a label (gpusharing / mps) with a value of "true", and the GPU of this type of node runs in the parallel running mode; the second type is the node with MPS disabled. This type of node does not have the "gpusharing / mps" label, or the value of this label is "false".Hijacking Scheduler: If this GPU node runs in parallel mode, the hijacking scheduler is only responsible for the allocation and limitation of video memory; if this GPU node runs in time-sharing scheduling mode, the hijacking scheduler is responsible for allocating time slices to Pods and limiting video memory allocation. Based on CUDA hijacking and MPS technology, this application can schedule multiple Pods to the same GPU, thereby achieving the goal of high GPU utilization. Moreover, the video memory of multiple Pods on the same GPU is isolated. Users can choose to execute multiple Pods on the same GPU in parallel to fully utilize the computing cores of the GPU, or multiple Pods can be scheduled in time-sharing and executed alternately. They can also choose to have a service exclusive to the GPU or share the GPU with other services.
[0059] The CUDA hijacking process of the Pod in this application is as Figure 4 shown. It is judged whether there is environment variable information MPS in the annotation information in the service creation information. If there is environment variable information MPS in the annotation information in the service creation information, it is judged whether the GPU sharing value in the annotation information is the same as the service sharing value obtained in advance. If the GPU sharing value in the annotation information is the same as the service sharing value obtained in advance, it is determined that the GPU sharing method is a service exclusive to the GPU. Then, from all GPU nodes, a GPU node with a GPU node label value of true and a GPU sharing method of a service exclusive to the GPU is selected as the target GPU node. The GPU running mode of the target GPU node is parallel mode. The environment variable information in the Pod is sent to the CUDA in the target GPU node so that the CUDA can perform GPU sharing calculations on the environment variable information using a preset calculation function and video memory application and release functions. If there is no environment variable information MPS in the annotation information in the service creation information, the GPU running mode is time-sharing scheduling mode. Then, it is determined that the GPU sharing method is a service sharing the GPU. Then, from all GPU nodes, a GPU node with a label value of false and a GPU sharing method of a service sharing the GPU is selected as the target GPU node. The CUDA determines the time ratio information from the environment variable information and performs GPU sharing calculations on the time ratio information using the video memory application and release functions. In addition, the present invention can be used not only on NVIDIA GPUs but also on other heterogeneous chips with a complete ecosystem.
[0060] In this embodiment, service creation information is obtained, a Pod is created based on the service creation information, and environment variable information is added to the Pod; the GPU operation mode and the GPU sharing mode are determined based on the service creation message; a target GPU node is determined according to the GPU sharing mode and the GPU operation mode, and the environment variable information in the Pod is sent to the target GPU node so that the target GPU node performs GPU sharing calculation on the environment variable information based on the GPU operation mode. Based on CUDA hijacking and MPS technology, this application can schedule multiple Pods to the same GPU, thereby achieving the purpose of high GPU utilization. Moreover, the video memories of multiple Pods on the same GPU are isolated. Users can choose to execute multiple Pods on the same GPU in parallel to make full use of the computing cores of the GPU, or schedule multiple Pods time-sharing and execute them alternately. They can also choose for the service to exclusively occupy the GPU or share the GPU with other services.
[0061] See Figure 5 As shown, an embodiment of the present invention discloses a GPU sharing device, which may specifically include:
[0062] A Pod creation module 11, configured to obtain service creation information, create a Pod based on the service creation information, and add environment variable information to the Pod;
[0063] A GPU operation mode determination module 12, configured to determine the GPU operation mode and the GPU sharing mode based on the service creation message;
[0064] A GPU sharing calculation module 13, configured to determine a target GPU node according to the GPU sharing mode and the GPU operation mode, and send the environment variable information in the Pod to the target GPU node so that the target GPU node performs GPU sharing calculation on the environment variable information based on the GPU operation mode.
[0065] In this embodiment, service creation information is obtained, a Pod is created based on the service creation information, and environment variable information is added to the Pod; the GPU operation mode and the GPU sharing mode are determined based on the service creation message; a target GPU node is determined according to the GPU sharing mode and the GPU operation mode, and the environment variable information in the Pod is sent to the target GPU node so that the target GPU node performs GPU sharing calculation on the environment variable information based on the GPU operation mode. Based on CUDA hijacking and MPS technology, this application can schedule multiple Pods to the same GPU, thereby achieving the purpose of high GPU utilization. Moreover, the video memory of multiple Pods on the same GPU is isolated. Users can choose to execute multiple Pods on the same GPU in parallel, making full use of the computing cores of the GPU, or schedule multiple Pods at different times and execute them alternately. They can also choose for the service to exclusively occupy the GPU or share the GPU with other services.
[0066] This application also includes an architecture diagram, and the main modules are introduced as follows: Master Node: The management node of the Kubernetes cluster, including a controller module and a scheduler module; Controller: Creates corresponding Pods and other resources for the services created by users, and injects the following environment variables into the containers of the Pods: MPS, which identifies whether the service uses MPS; Scheduler: Schedules the Pods corresponding to the services to a certain GPU and creates corresponding vGPUs; GPU Node: A computing node in the Kubernetes cluster with a GPU installed. There is a node agent running on such nodes, and the node agent includes the following sub-modules: Configuration Client: Responsible for registering the GPU resources of this node with the scheduler, and at the same time writing the list information of the Pods running on the GPU into a file: Hijack Scheduler: Reads the list information of the Pods written by the configuration client and completes the following functions: In the time-sharing scheduling mode, it is responsible for allocating time slices to the Pods and restricting the video memory allocation; In the parallel running mode, it is responsible for restricting the video memory allocation. When the scheduler schedules a newly created Pod, it will consider the following two conditions: Decide whether to schedule this Pod to a certain GPU on the GPU node in the parallel running mode or to a certain GPU on the time-sharing scheduling GPU node according to the annotation (gpusharing / mps) of this Pod; Decide whether the service of this Pod exclusively occupies a GPU or shares the same GPU with other services according to the annotation (gpusharing / group) of this Pod. If this annotation is not set in the yaml file of the service, then this service will share the same GPU with other services; If this annotation is set, then only the services with the same value of this annotation will share the same GPU; From the above explanation, it can be seen that if we want a service to exclusively occupy a GPU, then only need to make its "gpusharing / group" annotation unique (the meaning of exclusive occupation is: Suppose service 1 has 2 Pods, denoted as Pod 1 and Pod 2. Given that Pod 1 is scheduled to GPU node 1, then only Pods of service 1 can be on GPU 1, and Pods of other services cannot be scheduled to GPU node 1. That is to say, in the future, only Pod 2 may be scheduled to GPU node 1). Only the GPU nodes that meet both of the above two conditions will become the target GPU nodes. There are two types of GPU nodes in the cluster. The first type is the nodes with MPS enabled. These nodes will have a label (gpusharing / mps) with a value of "true", and the GPUs of these nodes run in the parallel running mode; The second type is the nodes with MPS disabled. These nodes do not have the "gpusharing / mps" label, or the value of this label is "false".Hijacking Scheduler: If this GPU node runs in the parallel running mode, then the hijacking scheduler is only responsible for the allocation and limitation of video memory; if this GPU node runs in the time-sharing scheduling mode, then the hijacking scheduler is responsible for allocating time slices to Pods and limiting the video memory allocation. Based on CUDA hijacking and MPS technology, this application can achieve scheduling multiple Pods to the same GPU, thereby achieving the purpose of high GPU utilization, and the video memories of multiple Pods on the same GPU are isolated. Users can choose to execute multiple Pods on the same GPU in parallel, making full use of the computing cores of the GPU, or multiple Pods can be scheduled in a time-sharing manner and executed alternately. They can also choose to have a service exclusively occupy the GPU or share the GPU with other services.
[0067] In some specific embodiments, the Pod creation module 11 may specifically include:
[0068] A detection module, configured to detect whether the creation of a local service is triggered;
[0069] An information acquisition module, configured to acquire service creation information if the creation of a local service is triggered.
[0070] In some specific embodiments, the Pod creation module 11 may specifically include:
[0071] A Pod creation module, configured to create the same number of Pods as the number of replicas based on the number of replicas in the service creation information;
[0072] An environment variable information adding module, configured to add environment variable information to the Pod based on the annotation information in the service creation information.
[0073] In some specific embodiments, the GPU operation mode determination module 12 may specifically include:
[0074] A GPU operation mode determination module, configured to determine the GPU operation mode based on the annotation information in the service creation message; wherein, the GPU operation mode includes a parallel running mode and a time-sharing scheduling mode;
[0075] A GPU sharing mode determination module, configured to determine the GPU sharing mode based on the annotation information in the service creation message; wherein, the GPU sharing mode includes a service exclusively occupying the GPU and a service sharing the GPU.
[0076] In some specific embodiments, the GPU operation mode determination module 12 may specifically include:
[0077] The first target GPU node determination module is configured to, if the GPU operation mode is the parallel operation mode, determine whether the GPU sharing value in the annotation information is the same as the service sharing value obtained in advance. If the GPU sharing value in the annotation information is the same as the service sharing value obtained in advance, determine that the GPU sharing mode is service exclusive GPU, and then screen out from all GPU nodes the GPU nodes with the GPU node label value being true and the GPU sharing mode being service exclusive GPU as the target GPU nodes;
[0078] The second target GPU node determination module is configured to, if the GPU operation mode is the time-sharing scheduling mode, determine that the GPU sharing mode is service shared GPU, and then screen out from all GPU nodes the GPU nodes with the label value being false and the GPU sharing mode being service shared GPU as the target GPU nodes.
[0079] In some specific embodiments, the GPU sharing calculation module 13 may specifically include:
[0080] The environment variable information sending module is configured to send the environment variable information in the Pod to the CUDA in the target GPU node, so that the CUDA in the target GPU node performs GPU sharing calculation on the environment variable information based on the GPU operation mode.
[0081] In some specific embodiments, the GPU sharing calculation module 13 may specifically include:
[0082] The first GPU sharing calculation module is configured to, if the GPU operation mode is the parallel operation mode, the CUDA performs GPU sharing calculation on the environment variable information by using a preset calculation function and a video memory application and release function;
[0083] The second GPU sharing calculation module is configured to, if the GPU operation mode is the time-sharing scheduling mode, the CUDA determines the time ratio information from the service creation information, and performs GPU sharing calculation on the time ratio information by using the calculation function.
[0084] Figure 6 This is a schematic structural diagram of an electronic device provided by an embodiment of the present application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. Among them, the memory 22 is used to store a computer program, and the computer program is loaded and executed by the processor 21 to implement the relevant steps in the GPU sharing method executed by the electronic device disclosed in any of the foregoing embodiments.
[0085] In this embodiment, the power supply 23 is used to provide operating voltages for the various hardware devices on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows can be any communication protocol applicable to the technical solution of this application, and no specific limitation is imposed on it here; the input / output interface 25 is used to obtain external input data or output data to the outside, and its specific interface type can be selected according to specific application needs, and no specific limitation is made here.
[0086] In addition, the memory 22, as a carrier for resource storage, can be a read-only memory, a random access memory, a magnetic disk, an optical disc, etc., and the resources stored thereon include an operating system 221, a computer program 222, data 223, etc., and the storage method can be transient storage or permanent storage.
[0087] Among them, the operating system 221 is used to manage and control the various hardware devices and the computer program 222 on the electronic device 20 to enable the processor 21 to perform operations and processing on the data 223 in the memory 22, and it can be Windows, Unix, Linux, etc. In addition to the computer program that can be used to complete the GPU sharing method executed by the electronic device 20 disclosed in any of the foregoing embodiments, the computer program 222 can further include computer programs that can be used to complete other specific tasks. In addition to the data transmitted by external devices received by the GPU sharing device, the data 223 can also include data collected by its own input / output interface 25, etc.
[0088] The steps of the method or algorithm described in combination with the embodiments disclosed in this document can be implemented directly by hardware, a software module executed by a processor, or a combination of the two. The software module can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the technical field.
[0089] Furthermore, the embodiments of this application also disclose a computer-readable storage medium, in which a computer program is stored, and when the computer program is loaded and executed by a processor, the steps of the GPU sharing method disclosed in any of the foregoing embodiments are implemented.
[0090] Finally, it should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising said element.
[0091] The above has introduced in detail a GPU sharing method, apparatus, device and storage medium provided by the present invention. Specific examples are used in this text to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A GPU sharing method, characterized in that, Including: Obtain service creation information, create a Pod based on the service creation information, and add environment variable information to the Pod; Determine the GPU operation mode and GPU sharing mode based on the service creation information; Determine a target GPU node according to the GPU sharing mode and the GPU operation mode, and send the environment variable information in the Pod to the target GPU node, so that the target GPU node performs GPU sharing calculation on the environment variable information based on the GPU operation mode; Among them, create Pods with the same number as the number of replicas based on the number of replicas in the service creation information; add environment variable information to the Pods based on the annotation information in the service creation information; determine the GPU operation mode based on the annotation information in the service creation information; determine the GPU sharing mode based on the annotation information in the service creation information; the GPU operation mode includes a parallel operation mode and a time-sharing scheduling mode; the GPU sharing mode includes service-exclusive GPU and service-shared GPU; Sending the environment variable information in the Pod to the target GPU node so that the target GPU node performs GPU sharing calculation on the environment variable information based on the GPU operation mode includes: sending the environment variable information in the Pod to the CUDA in the target GPU node. If the GPU operation mode is the parallel operation mode, the CUDA uses a preset calculation function and video memory application and release functions to perform GPU sharing calculation on the environment variable information; if the GPU operation mode is the time-sharing scheduling mode, the CUDA determines the time ratio information from the service creation information and performs GPU sharing calculation on the time ratio information using the calculation function.
2. The GPU sharing method according to claim 1, wherein The obtaining of the service creation information includes: Detect whether the local service creation is triggered; If the local service creation is triggered, obtain the service creation information.
3. The GPU sharing method according to claim 1, wherein The determining of the target GPU node according to the GPU sharing mode and the GPU operation mode includes: If the GPU operation mode is the parallel operation mode, determine whether the GPU sharing value in the annotation information is the same as the pre-obtained service sharing value. If the GPU sharing value in the annotation information is the same as the pre-obtained service sharing value, determine that the GPU sharing mode is service-exclusive GPU, and then screen out the GPU nodes with the GPU node label value of true and the GPU sharing mode of service-exclusive GPU from all GPU nodes as the target GPU nodes; If the GPU operation mode is the time-sharing scheduling mode, determine that the GPU sharing mode is service-shared GPU, and then screen out the GPU nodes with the label value of false and the GPU sharing mode of service-shared GPU from all GPU nodes as the target GPU nodes.
4. A GPU sharing device, characterized in that, Including: A Pod creation module, configured to obtain service creation information, create a Pod based on the service creation information, and add environment variable information to the Pod; A GPU operation mode determination module, configured to determine the GPU operation mode and the GPU sharing mode based on the service creation information; A GPU shared computing module, configured to determine a target GPU node according to the GPU sharing method and the GPU operation mode, and send the environment variable information in the Pod to the target GPU node, so that the target GPU node performs GPU shared computing on the environment variable information based on the GPU operation mode; Among them, create Pods with the same number as the number of replicas based on the number of replicas in the service creation information; add environment variable information to the Pods based on the annotation information in the service creation information; determine the GPU operation mode based on the annotation information in the service creation information; determine the GPU sharing method based on the annotation information in the service creation information; the GPU operation mode includes a parallel operation mode and a time-sharing scheduling mode; the GPU sharing method includes service exclusive GPU and service shared GPU; Sending the environment variable information in the Pod to the target GPU node, so that the target GPU node performs GPU shared computing on the environment variable information based on the GPU operation mode, includes: sending the environment variable information in the Pod to the CUDA in the target GPU node, if the GPU operation mode is the parallel operation mode, then CUDA performs GPU shared computing on the environment variable information by using a preset computing function and a video memory application and release function; if the GPU operation mode is the time-sharing scheduling mode, then CUDA determines the time ratio information from the service creation information, and performs GPU shared computing on the time ratio information by using the computing function.
5. An electronic device, characterized in that, Includes: A memory, configured to store a computer program; A processor, configured to execute the computer program to implement the GPU sharing method according to any one of claims 1 to 3.
6. A computer-readable storage medium, characterized in that, For storing a computer program; wherein, when the computer program is executed by the processor, the GPU sharing method according to any one of claims 1 to 3 is implemented.
Citation Information
Patent Citations
GPU sharing scheduling and single-machine multi-card method, system and device
CN111475303A
Computing device sharing method and apparatus based on kubernetes, and device and storage medium
WO2022062650A1