Implementation method and device for assigning GPU to create load in Kubernetes and storage medium
By introducing custom plug-ins and Volcano scheduler to filter GPU resources in Kubernetes, and combining OCI Hook to achieve refined GPU management, the problem of inflexible allocation of GPU resources in the existing technology is solved, and the direct binding and precise allocation of containers and GPUs are realized.
Patent Information
- Application Number
- CN202510547726.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-08-08
AI Technical Summary
The existing Kubernetes platform cannot granularly control the allocation of GPU resources, resulting in unclear binding relationship between containers and GPUs, and it is impossible to flexibly allocate specified GPUs for different workloads.
During the Kubernetes scheduling stage, the GPU resources that meet the conditions are filtered out through custom plug-ins, and the specified GPU is mapped to the container through kubelet, and the GPU is refined and direct binding is achieved using Volcano scheduler and OCI Hook.
It realizes the precise allocation of specified GPU cards during Pod scheduling, adapts to the needs of different scenarios, and establishes a clear container-GPU relationship at the Kubernetes level.
Smart Images

Figure CN120448119A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of containerized platform deployment, and in particular to an implementation method, device, and storage medium for supporting load creation by a specified GPU in Kubernetes. Background Art
[0002] Current mainstream containerization platforms, such as Kubernetes (K8s), do not, by default, support the granularity of specifying which GPUs (Graphics Processing Units) on a node are used to create workloads. While Kubernetes provides basic GPU resource scheduling capabilities, there are still some limitations in fine-grained GPU management.
[0003] Kubernetes has introduced support for GPU resources since version 1.10, primarily for scheduling computational loads that require GPUs. By integrating with GPU drivers (such as the NVIDIA GPU driver) and Kubernetes' node-level plug-ins, Kubernetes can identify available GPUs on a node and allocate them as part of scheduling resources. Kubernetes allows users to request GPU resources in a Pod (the smallest scheduling unit in Kubernetes, consisting of one or more containers) definition, as shown in the following example:
[0004] resources:
[0005] limits:
[0006] nvidia.com / gpu:1#Request 1 GPU
[0007] This mechanism allows Kubernetes to identify GPUs and schedule them to nodes with GPUs. However, this default mechanism in Kubernetes does not provide fine-grained control over the allocation of GPU resources.
[0008] In the Kubernetes platform, existing GPU allocation technology solutions are primarily implemented through the interaction between the kubelet (a key component of the Kubernetes framework responsible for controlling and coordinating pods and nodes) and the device plugin (an interface reserved by Kubernetes to enable containers to use a wider range of computing resources). During this process, the allocate interface returns an available GPU card through the environment variable NVIDIA_VISIBLE_DEVICES. The underlying container runtime uses this variable to set the GPU binding relationship between the host and the container. However, this process has certain limitations, particularly regarding the binding relationship between containers and GPUs.
[0009] The key steps of the current GPU allocation technology solution include:
[0010] 1. Kubernetes schedules GPU load: When users request GPU resources in the Pod definition, the Kubernetes scheduler recognizes these resource requests and schedules the load to nodes with available GPUs.
[0011] 2. Kubelet interacts with the Device Plugin: The kubelet on each node interacts with the Kubernetes device plugin (such as the NVIDIA Device Plugin). The device plugin provides a list of all GPU devices on the node and allows the kubelet to allocate available GPUs by calling the allocate interface.
[0012] 3. Allocate interface returns GPU information: When kubelet calls the allocate interface, the device plugin selects one or more available GPUs based on its internal logic and returns the ID of the selected GPU to kubelet through the environment variable NVIDIA_VISIBLE_DEVICES.
[0013] 4. NVIDIA_VISIBLE_DEVICES environment variable: The GPU ID returned by the allocate API is stored in the NVIDIA_VISIBLE_DEVICES environment variable, which tells the underlying container runtime which GPUs to use. When the container runtime starts a container, it establishes a GPU device mapping between the host and the container based on the GPU ID specified in NVIDIA_VISIBLE_DEVICES.
[0014] The above solution has the following defects:
[0015] The device plugin selects one or more available GPUs based on its internal logic. This allocation is usually random or based on a simple policy, and there is no way to control the specific GPU or GPUs to be used. This can lead to the following:
[0016] (1) Binding between containers and GPUs can only be achieved through environment variables and device mapping during the underlying container runtime. There is no direct relationship between containers and GPUs at the Kubernetes level.
[0017] (2) It is not possible to flexibly assign specific GPUs to different workloads, which cannot meet user needs, such as users requiring exclusive use of a specific GPU. Summary of the Invention
[0018] In response to the defects in the existing technology, the present invention solves the technical problem of how to create a load based on a specified GPU in Kubernetes.
[0019] To achieve the above objectives, in the first aspect, an embodiment of the present application provides an implementation method for creating a load by specifying a GPU in Kubernetes, the method comprising the following steps: obtaining the GPU information that needs to be specified, filtering out the nodes and GPU IDs corresponding to the GPU information in the specified screening phase during Kubernetes scheduling, and sending them to the kubelet of the node corresponding to the GPU information; starting a container through the kubelet and mapping the GPU corresponding to the GPU ID to the container.
[0020] In combination with the first aspect, in one embodiment, the specified screening stage is the Predicates stage; the process of screening out nodes and GPU IDs corresponding to GPU information in the specified screening stage during Kubernetes scheduling includes: adding a custom plug-in to the Volcano scheduler, and using the custom plug-in to screen out nodes and GPU IDs corresponding to the GPU information from all nodes and GPUs.
[0021] In conjunction with the first aspect, in one embodiment, the method for screening the nodes and GPU IDs corresponding to the GPU information is:
[0022] When the node and GPU ID specified in the GPU information are not empty and the corresponding node and GPU are available, the node and GPU ID corresponding to the GPU information are returned;
[0023] When the node and GPU ID specified in the GPU information are not empty and the corresponding GPU is unavailable, the message "No GPU ID available on the node corresponding to the GPU information" is returned.
[0024] When the node specified in the GPU information is not empty and the GPU ID is empty, the node corresponding to the GPU information and the score list of all available GPUs in the node are returned;
[0025] When the node and GPU ID specified in the GPU information are both empty, a score list of all available nodes and a score list of all available GPU IDs are returned.
[0026] In combination with the first aspect, in one embodiment, the process of sending to the kubelet of the node corresponding to the GPU information includes: when there is a scoring list of nodes and / or GPUs, selecting the node and / or GPU ID with the highest score in the scoring list of nodes and / or GPUs and sending it to the kubelet.
[0027] In combination with the first aspect, in one embodiment, the process of starting a container through kubelet and mapping the GPU corresponding to the GPUID to the container includes: kubelet initiates a container creation request to OCI Hook including a node and GPU ID corresponding to the GPU information.
[0028] In combination with the first aspect, in one embodiment, before starting the container through kubelet and mapping the GPU corresponding to the GPU ID to the container, the following steps are also included: kubelet remotely calls the Device Plugin through the allocate interface as a transmission channel.
[0029] In combination with the first aspect, in one implementation, the process of obtaining the GPU information that needs to be specified includes: specifying the GPU information in the annotation of the pod.
[0030] In combination with the first aspect, in one embodiment, the execution process of obtaining the GPU information that needs to be specified includes: the client sends a POD creation request containing GPU information to the Kubernetes server, and the Kubernetes server returns a creation success message to the client when it is normal.
[0031] In a second aspect, an embodiment of the present application provides an implementation device for creating a load on a specified GPU in Kubernetes, wherein the implementation device for creating a load on a specified GPU in Kubernetes includes a processor, a memory, and an implementation program for creating a load on a specified GPU in Kubernetes stored in the memory and executable by the processor, wherein when the implementation program for creating a load on a specified GPU in Kubernetes is executed by the processor, the steps of the method provided in the first aspect are implemented.
[0032] In a third aspect, an embodiment of the present application provides a computer-readable storage medium, on which is stored an implementation program for specifying a GPU to create a load in Kubernetes. When the implementation program for specifying a GPU to create a load in Kubernetes is executed, the steps of the method provided in the first aspect are implemented.
[0033] Compared with the prior art, the advantages of the present invention are:
[0034] The present invention can screen out qualified (need to be specified) GPU resources during the Kubernetes scheduling phase. This scheduling process combines refined management of GPU resources and can accurately allocate specified GPU cards during Pod scheduling. This not only allows for the allocation of specified GPUs for different workloads to meet the needs of different scenarios, but also establishes a direct relationship between containers and GPUs at the Kubernetes level. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0036] Figure 1 This is a flow chart of a method for implementing a load creation for a specified GPU in Kubernetes in an embodiment of the present invention;
[0037] Figure 2 A schematic diagram of the hardware structure of a device for implementing load creation for a specified GPU in Kubernetes involved in the embodiment of the present application. DETAILED DESCRIPTION
[0038] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0039] The flowcharts shown in the accompanying drawings are for illustrative purposes only and do not necessarily include all contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps may be decomposed, combined, or partially merged, so the actual execution order may vary depending on the actual situation.
[0040] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.
[0041] In the first aspect, an embodiment of the present application provides an implementation method for specifying GPU creation load in Kubernetes, the steps of which include: obtaining the GPU information that needs to be specified (including at least the GPU ID and the node location of the GPU), filtering out the nodes and GPU IDs corresponding to the GPU information in the specified screening phase during Kubernetes scheduling, and sending them to the kubelet of the node corresponding to the GPU information; starting the container through the kubelet and mapping the GPU corresponding to the GPU ID to the container.
[0042] It can be seen from this that the present invention can screen out qualified (need to be specified) GPU resources during the scheduling stage of Kubernetes. This scheduling process is combined with the refined management of GPU resources and can accurately allocate specified GPU cards during Pod scheduling. It can not only allocate specified GPUs for different workloads to meet the needs of different scenarios, but also establish a direct relationship between containers and GPUs at the Kubernetes level.
[0043] In one embodiment, the process of obtaining the GPU information that needs to be specified in the above method includes: specifying the GPU information in the annotation of the pod.
[0044] On this basis, the execution process for obtaining the required GPU information can be as follows: the user's client sends a POD creation request containing GPU information to the Kubernetes server, and the Kubernetes server returns a creation success message to the client when normal.
[0045] In one embodiment, the specified screening stage in the above method is the Predicates stage, which is an important step in Kubernetes scheduling decisions. On this basis, the process of screening out nodes and GPU IDs corresponding to GPU information in the specified screening stage during Kubernetes scheduling in the above method includes: adding a custom plug-in to the Volcano scheduler, and filtering out nodes and GPU IDs corresponding to GPU information from all nodes and GPUs through the custom plug-in. Volcano is a commonly used batch job scheduling framework on Kubernetes, which is good at scheduling complex resource-intensive tasks (such as big data analysis, deep learning, etc.). It has more flexible scheduling expansion capabilities and supports plug-in scheduling decisions. On this basis, by adding only one custom plug-in to Volcano, GPUs can be screened according to GPU usage and POD GPU requirements without significantly changing the code, thereby realizing the effect of maintaining the usage status of each GPU card on the node through the plug-in to ensure that the POD can be allocated to a specific GPU according to demand.
[0046] In one embodiment, the method for screening nodes and GPU IDs corresponding to GPU information in the above method is:
[0047] When the node and GPU ID specified in the GPU information are not empty and the corresponding node and GPU are available, the node and GPU ID corresponding to the GPU information are returned;
[0048] If the node and GPU ID specified in the GPU information are not empty and the corresponding GPU is unavailable, the system returns "No GPU ID available on the node corresponding to the GPU information". It also returns a list of scores of all available nodes and a list of scores of all available GPU IDs (the score calculation method is the existing technology) for subsequent Kubernetes server selection.
[0049] When the node specified in the GPU information is not empty and the GPU ID is empty, the node corresponding to the GPU information and the score list of all available GPUs in the node are returned;
[0050] When the node and GPU ID specified in the GPU information are both empty, a score list of all available nodes and a score list of all available GPU IDs are returned.
[0051] On this basis, the process of sending to the kubelet of the node corresponding to the GPU information in the above method includes: when there is a scoring list of nodes and / or GPUs, selecting the node and / or GPU ID with the highest score in the scoring list of nodes and / or GPUs and sending it to the kubelet; of course, if the above return is the specified node and GPU ID, it will be sent directly.
[0052] In one embodiment, before the above method starts the container through kubelet and maps the GPU corresponding to the GPU ID to the container, it also includes the following steps: kubelet remotely calls DevicePlugin (a standardized resource extension mechanism in Kubernetes, which aims to integrate special hardware as schedulable resources into the resource management framework of Kubernetes) through the allocate interface as a transmission channel. The execution method can be that kubelet sends an RPC (remote call) request containing the specified node and GPU ID to Device Plugin, and Device Plugin returns the RPC request success to kubelet.
[0053] From this we can see that:
[0054] (1) In the present invention, the allocate interface is no longer responsible for the actual allocation of GPU resources, but is just a channel that does not interfere with the selection of GPU cards (avoiding the random allocation logic of the allocat interface). The selection of GPUs has been completed by the custom plug-in of the Volcano scheduler.
[0055] (2) When the GPU selection has been completed by the custom plug-in of the Volcano scheduler, the container can be started directly and the GPU can be mapped. Therefore, remotely calling the Device Plugin through kubelet is a redundant step. However, since remotely calling the Device Plugin through kubelet is a native step of Kubernetes and is upstream code, it is difficult to change the code here. Therefore, a "redundant" step is used to make it compatible with the Kubernetes workflow.
[0056] In one embodiment, the process of starting a container through kubelet and mapping the GPU corresponding to the GPUID to the container in the above method includes: kubelet initiates a container creation request including the node and GPU ID corresponding to the GPU information to OCI Hook (which comes with Kubernetes and is part of the OCI standard. It aims to provide a mechanism for users or developers to insert custom hook functions at different life cycle stages of the container).
[0057] When starting the container, OCI Hook uses the --device parameter to expose the GPU corresponding to the GPU ID and map it to the container.
[0058] At the same time, the --device parameter is used to expose the specified GPU device to the container, enabling direct mapping of the GPU device. This step ensures that the binding relationship between the container and the GPU is clear and controllable, breaking the randomness of the existing NVIDIA_VISIBLE_DEVICES environment variable.
[0059] See below Figure 1 The specific interaction process of the above method is described through a specific embodiment.
[0060] S1: The user enters GPU information in the GPU POD (a container in the Kubernetes cluster specifically used to run applications that require GPU acceleration). The GPU POD sends a POD creation request containing the GPU information to the KUBE-APIServer (Kubernetes server). If the KUBE-APIServer is functioning normally, it returns a creation success message to the client.
[0061] S2: KUBE-APIServer sends a filtering instruction to the Volcano scheduler. The Volcano scheduler filters in the Predicates stage (see above for the specific filtering method) to obtain the node and GPU ID corresponding to the GPU information and returns it to the KUBE-APIServer.
[0062] S3: KUBE-APIServer sends the node and GPU ID corresponding to the GPU information to Kubelet.
[0063] S4: Kubelet sends an RPC (remote procedure call) request containing the specified node and GPU ID to the Device Plugin through the allocate interface as a transmission channel. The Device Plugin returns the RPC request success to the kubelet.
[0064] S5: The kubelet initiates a container creation request to the OCI Hook, including the node and GPU ID corresponding to the GPU information. The OCI Hook obtains the Pod's annotations information through the annotations field in the hook state, which includes the GPU ID previously assigned by the Volcano scheduler. When starting the container, the OCI Hook uses the --device parameter to expose the GPU corresponding to the GPU ID and map it to the container; this allows the started container to directly use the assigned GPU.
[0065] In a second aspect, an embodiment of the present application provides an implementation device for creating a load on a specified GPU in Kubernetes. The implementation device for creating a load on a specified GPU in Kubernetes can be a personal computer (PC), a laptop, a server, or other device with data processing capabilities.
[0066] Reference Figure 2 , Figure 2 This is a hardware structure diagram of the device for implementing the creation of a load on a specified GPU in Kubernetes involved in the embodiment of the present application. In the embodiment of the present application, the device for implementing the creation of a load on a specified GPU in Kubernetes may include a processor, a memory, a communication interface, and a communication bus.
[0067] The communication bus may be of any type and is used to interconnect the processor, memory, and communication interface.
[0068] Communication interfaces include input / output (I / O) interfaces, physical interfaces, and logical interfaces. These interfaces interconnect devices within the device that creates a specific GPU workload in Kubernetes, as well as interfaces that connect the device to other devices (such as other computing devices or user devices). Physical interfaces can include Ethernet, fiber, or ATM interfaces; user devices can include displays and keyboards.
[0069] The memory can be various types of storage media, such as random access memory (RAM), read-only memory (ROM), non-volatile RAM (NVRAM), flash memory, optical storage, hard disk, programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), etc.
[0070] The processor may be a general-purpose processor, which may call the implementation program for specifying GPU creation load in Kubernetes stored in the memory, and execute the implementation method for specifying GPU creation load in Kubernetes provided in the embodiment of the present application. For example, the general-purpose processor may be a central processing unit (CPU). The method executed when the implementation program for specifying GPU creation load in Kubernetes is called may refer to the various embodiments of the implementation method for specifying GPU creation load in Kubernetes of the present application, which will not be repeated here.
[0071] Those skilled in the art will understand that Figure 2 The hardware structure shown in the figure does not constitute a limitation to the present application and may include more or fewer components than shown in the figure, or a combination of certain components, or a different arrangement of components.
[0072] In a third aspect, an embodiment of the present application also provides a computer-readable storage medium.
[0073] The computer-readable storage medium of the present application stores an implementation program for creating a load by specifying a GPU in Kubernetes. When the implementation program for creating a load by specifying a GPU in Kubernetes is executed by a processor, the steps of the implementation method for creating a load by specifying a GPU in Kubernetes as described above are implemented.
[0074] Among them, the method implemented when the implementation program of specifying GPU to create load in Kubernetes is executed can refer to the various embodiments of the implementation method of specifying GPU to create load in Kubernetes of this application, and will not be repeated here.
[0075] It should be noted that the serial numbers of the above-mentioned embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.
[0076] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, of course, it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes a number of instructions for enabling a terminal device to execute the methods described in each embodiment of the present application.
[0077] The terms "including" and "having" and any variations thereof in the specification and claims of this application and the above-mentioned drawings are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally includes steps or units that are not listed, or optionally includes other steps or units inherent to these processes, methods, products or devices. The terms "first", "second" and "third" are used to distinguish different objects, etc., and do not represent a sequence, nor do they limit the "first", "second" and "third" to different types.
[0078] In the description of the embodiments of this application, the words "exemplary," "for example," or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary," "for example," or "for example" in the embodiments of this application should not be construed as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary," "for example," or "for example" is intended to present the relevant concepts in a concrete manner.
[0079] In the description of the embodiments of the present application, unless otherwise specified, “ / ” means or, for example, A / B can mean A or B; “and / or” in the text is merely a description of the association relationship of associated objects, indicating that three relationships may exist, for example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, in the description of the embodiments of the present application, “multiple” refers to two or more than two.
[0080] In some processes described in the embodiments of the present application, multiple operations or steps are included that appear in a specific order. However, it should be understood that these operations or steps may not be performed in the order in which they appear in the embodiments of the present application or may be performed in parallel. The sequence numbers of the operations are only used to distinguish between different operations, and the sequence numbers themselves do not represent any order of execution. In addition, these processes may include more or fewer operations, and these operations or steps may be performed in sequence or in parallel, and these operations or steps may be combined.
[0081] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, of course, it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes a number of instructions for enabling a terminal device to execute the methods described in each embodiment of the present application.
[0082] The above are only specific implementations of the embodiments of the present invention, but the scope of protection of the embodiments of the present invention is not limited to them. Any person skilled in the art can easily conceive of various equivalent modifications or replacements within the technical scope disclosed in the embodiments of the present invention, and such modifications or replacements should be included in the scope of protection of the embodiments of the present invention. Therefore, the scope of protection of the embodiments of the present invention should be based on the scope of protection of the claims.
Claims
1. A method for implementing a load creation on a specified GPU in Kubernetes, characterized in that: The method includes the following steps: obtaining the GPU information that needs to be specified, filtering out the node and GPU ID corresponding to the GPU information in the specified filtering phase during Kubernetes scheduling, and sending it to the kubelet of the node corresponding to the GPU information; starting the container through the kubelet and mapping the GPU corresponding to the GPU ID to the container.
2. The method for implementing a load creation on a designated GPU in Kubernetes according to claim 1, wherein: The designated screening stage is the Predicates stage; The process of filtering out nodes and GPU IDs corresponding to GPU information in the specified filtering phase during Kubernetes scheduling includes: adding a custom plug-in to the Volcano scheduler, and using the custom plug-in to filter out nodes and GPU IDs corresponding to GPU information from all nodes and GPUs.
3. The method for implementing a load creation by specifying a GPU in Kubernetes according to claim 1, wherein: The method for screening the nodes and GPU IDs corresponding to the GPU information is as follows: When the node and GPU ID specified in the GPU information are not empty and the corresponding node and GPU are available, the node and GPU ID corresponding to the GPU information are returned; When the node and GPU ID specified in the GPU information are not empty and the corresponding GPU is unavailable, the message "No GPU ID available on the node corresponding to the GPU information" is returned. When the node specified in the GPU information is not empty and the GPU ID is empty, the node corresponding to the GPU information and the score list of all available GPUs in the node are returned; When the node and GPU ID specified in the GPU information are both empty, a score list of all available nodes and a score list of all available GPU IDs are returned.
4. The method for implementing a load creation by specifying a GPU in Kubernetes according to claim 3, wherein: The process of sending to the kubelet of the node corresponding to the GPU information includes: when there is a scoring list of nodes and / or GPUs, selecting the node and / or GPU ID with the highest score in the scoring list of nodes and / or GPUs and sending it to the kubelet.
5. The method for implementing a load creation by specifying a GPU in Kubernetes according to claim 1, wherein: The process of starting a container through kubelet and mapping the GPU corresponding to the GPUID to the container includes: kubelet initiates a container creation request including the node and GPU ID corresponding to the GPU information to OCI Hook.
6. The method for implementing a load creation by specifying a GPU in Kubernetes according to claim 1, wherein: Before starting the container through kubelet and mapping the GPU corresponding to the GPU ID to the container, the following steps are also included: kubelet remotely calls the Device Plugin through the allocate interface as a transmission channel.
7. The method for implementing load creation by specifying a GPU in Kubernetes according to any one of claims 1 to 6, wherein: The process of obtaining the GPU information that needs to be specified includes: specifying the GPU information in the pod's annotation.
8. The method for implementing a load creation on a designated GPU in Kubernetes according to claim 7, wherein: The execution process of obtaining the GPU information that needs to be specified includes: the client sends a POD creation request containing GPU information to the Kubernetes server, and the Kubernetes server returns a creation success message to the client when it is normal.
9. A device for implementing load creation on a specified GPU in Kubernetes, characterized in that: The implementation device for specifying GPU creation load in KUBERNETES includes a processor, a memory, and an implementation program for specifying GPU creation load in KUBERNETES stored in the memory and executable by the processor, wherein when the implementation program for specifying GPU creation load in KUBERNETES is executed by the processor, the steps of the implementation method for specifying GPU creation load in KUBERNETES as described in any one of claims 1 to 8 are implemented.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores an implementation program for specifying a GPU to create a load in Kubernetes. When the implementation program for specifying a GPU to create a load in Kubernetes is executed, the steps of the implementation method for specifying a GPU to create a load in Kubernetes as described in any one of claims 1 to 8 are implemented.
Citation Information
Cited By
Resource management method, electronic equipment and storage medium
CN121037187A
Implementation method and device for creating load by specified GPU, equipment and storage medium
CN121233349A
Method, device and storage medium for creating a load implementation method by designating a GPU
CN121233349B