Multi-resource mixed running cloud native elastic GPU virtualization method and system and storage medium thereof
Through the multi-resource hybrid running method of fine-grained time slice segmentation and kernel-state driver monitoring, the low resource utilization and ecological compatibility of GPU virtualization technology in the cloud-native environment is solved, efficient and flexible scheduling and isolation of GPU resources are achieved, and overall efficiency and cost-effectiveness are improved.
Patent Information
- Application Number
- CN202510334680.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-07-08
AI Technical Summary
The existing GPU virtualization technology has problems such as low resource utilization, lack of unified standards in closed ecosystems, large performance overhead, and inability to adapt flexibly in a flexible cloud-native environment, especially when sharing GPU resources, it is difficult to achieve efficient utilization.
The cloud-native elastic GPU virtualization method with multi-resource hybrid running is adopted. Through fine-grained time slice segmentation and kernel-state driver monitoring, the running status of GPU tasks is dynamically adjusted, and combined with the container SLA guarantee mechanism, the isolation and flexible scheduling of GPU computing power and video memory are achieved.
It improves the utilization rate and overall efficiency of GPU resources, reduces the cost of use, and is compatible with the cloud-native ecosystem, ensuring the execution experience and fault isolation capabilities of AI tasks.
Smart Images

Figure CN120276840A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of GPU virtualization, and specifically to a cloud-native elastic GPU virtualization method, system and storage medium for mixed running of multiple resources. Background Technique
[0002] With the popularity of concepts such as large models and AGI, HPC has gradually split into two forms: distributed training and inference. Combining with the cloud-native direction that has developed the fastest in the past few years, the SLA of inference tasks is gradually being taken seriously. The demand for shared GPU cards is becoming increasingly mature, and GPU virtualization technology has become the infrastructure foundation for sharing GPUs in inference tasks. GPU virtualization solutions stem from the effective utilization and flexible adaptation of computing power. In a production environment, a single inference task often cannot fully utilize the resources of an entire card;
[0003] Moreover, the GPU ecosystem is closed and lacks a unified standard; GPU virtualization technology has a certain binding relationship with GPU manufacturers. For example, most CUDA-compatible GPU products use the CUDA hijacking scheme to achieve virtualization, and overload the CUDA driver API through the preload mechanism. The CUDA hijacking scheme has good system compatibility, but the disadvantage is that the performance overhead is relatively large. Also, in the scenarios of cloud-native or public clouds, the business usually cannot accept modifying the container image;
[0004] Based on business requirements such as performance and non-intrusiveness, we propose a cloud-native elastic GPU virtualization solution for mixed running of multiple resources Summary of the Invention
[0005] The purpose of the present invention is to provide a cloud-native elastic GPU virtualization method, system and storage medium for mixed running of multiple resources, so that the GPU computing power can be optimized in terms of supporting service orientation, flexible computing power scheduling, and effective utilization rate, while ensuring the cost and being well combined with the cloud-native ecosystem, so as to solve the problems raised in the above background technique.
[0006] To achieve the above purpose, the present invention provides the following technical solution: A cloud-native elastic GPU virtualization method for mixed running of multiple resources, including performing fine-grained time slicing on GPU computing resources, allocating each sliced time slice to the corresponding Pod for use, and monitoring the running status of GPU tasks in real time using the kernel-mode driver, intervening in the Pods whose running time and video memory usage exceed the allocated specifications, dynamically adjusting and selecting the designed GPU computing power scheduling strategy according to the usage scenario, and at the same time using the container SLA guarantee mechanism to sense the health of business containers and the GPU load situation, and performing dynamic node scheduling on the Pods.
[0007] Preferably, the fine-grained time slice division supports 1% granularity GPU division. The usage right of the GPU card is divided into a series of time slices and shared by the Pods on the k8sNode. If there is no GPU computing requirement for the current Pod within a time slice, this time slice will not be allocated to other Pods for use.
[0008] Preferably, the method for intervening in a Pod whose running time exceeds the allocated specification includes that the hypervisor marks the Pod whose running time of the CUDA program exceeds its allocated specification within the current time slice as blocked, and then deprives the Pod of its GPU usage right through the task scheduling subsystem of the Linux kernel.
[0009] Preferably, the method for intervening in a Pod whose video memory usage exceeds the allocated specification includes monitoring and summarizing the video memory usage in the Pod through the kernel driver module. When the video memory usage exceeds the Pod specification, a video memory allocation failure is directly returned.
[0010] Preferably, the scheduling policy includes:
[0011] Fixed computing power mode. The Pods on the Node use the computing power according to the specification, and the hypervisor strictly provides the computing power according to the Pod specification. Even if the current Pod has no GPU computing task, this time slice will not be allocated to other Pods for use;
[0012] Dynamic computing power mode. All the Pods on the Node evenly divide the idle computing power. When the total computing power specification of all Pods within the current scheduling period is less than 100%, or some Pods have no GPU computing task within the scheduling time slice, the hypervisor considers that there is idle computing power in the current scheduling period, and it will evenly allocate the remaining time slices to each Pod for use;
[0013] Minimum computing power mode. All the Pods on the Node will obtain at least the computing power resources declared by the Pod specification. When the total computing power specification of all Pods within the current scheduling period is less than 100%, or some Pods have no GPU computing task within the scheduling time slice, the hypervisor considers that there is idle computing power in the current scheduling period, and it will randomly allocate the remaining time slices to the Pods for use;
[0014] Priority scheduling computing power mode. Overbooking of computing power is provided on the Node. When the total computing power specification of all Pods within the current scheduling period is greater than 100%, high-priority tasks can preempt the time slices of low-priority tasks. This solution provides an additional interface in the cgroup directory to configure the Pod priority.
[0015] Preferably, the container SLA guarantee mechanism senses the health of business containers and the GPU load, and feeds them back to the apiserver through kubelet to adjust the cluster scheduling policy.
[0016] Preferably, the Pod dynamic node scheduling includes:
[0017] Scheduling based on the health status, where if a Pod is in poor health, attempt to reschedule it to other nodes;
[0018] Scheduling based on GPU load, where if the GPU load on a Node is too high, preferentially schedule new Pods to idle nodes;
[0019] Adjustment based on priority: For low-priority tasks, they can be actively evicted to release resources.
[0020] Preferably, monitoring the running status of GPU tasks includes that the hypervisor intercepts the event of a CUDA program submitting a task to the GPU through the kernel driver module. When initializing the kernel driver module, a hook function is registered to intercept the CUDA task submission operation. When the CUDA program calls the driver API to submit a task, the hook function will be triggered to capture relevant events, and the hook function records information such as the type of the task, parameters, and virtual machine ID.
[0021] To solve the above technical problems, the present invention also provides a cloud-native elastic GPU virtualization system for multi-resource mixed running, including:
[0022] A memory for storing computer programs;
[0023] A processor for executing the computer programs, and when the computer programs are executed by the processor, the steps of a cloud-native elastic GPU virtualization method for multi-resource mixed running as described in any one of the above are implemented.
[0024] To solve the above technical problems, the present invention also provides a readable storage medium with computer programs stored thereon,
[0025] When the computer programs are executed by a processor, the steps of a cloud-native elastic GPU virtualization method for multi-resource mixed running as described in any one of the above are implemented.
[0026] In summary, the beneficial effects of the present invention are:
[0027] 1. The present invention optimizes the GPU computing power in terms of supporting service - orientation, flexible computing power scheduling, and effective utilization rate by considering both cost and well - integrating with the cloud - native ecosystem. This solution abstracts the execution environment of AI tasks, provides virtual GPU products with specific computing power and video memory specifications. Each service running on the virtual GPU card believes that it exclusively occupies the entire card's resources. It also provides isolation technologies for computing power and video memory to ensure the usage experience and fault isolation ability of AI tasks.
[0028] 2. On the basis of achieving isolation between computing power and video memory and decoupling the user from the physical card, the present invention further optimizes the scenario of low overall card utilization rate by providing the ability to mix and run multiple tasks with elastic computing power; the computing power other than that allocated to the virtual GPU is allocated to all Pods on the Node according to certain rules, increasing the collaborative ability of elastic computing power during the window period with abundant resources, so as to improve the actual utilization rate of the GPU card, making it possible to mix and run tasks of multiple resource types, greatly improving resource efficiency, and reducing the usage cost of the GPU card through technological dividends.
[0029] The present invention also provides a cloud - native elastic GPU virtualization system for multi - resource mixed running and its storage medium, which has the above - mentioned beneficial effects and will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] In order to more clearly illustrate the technical solutions in the embodiments of the invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following - described drawings are only some embodiments of the invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0031] Figure 1 It is a schematic flowchart of a cloud - native elastic GPU virtualization method for multi - resource mixed running of the present invention;
[0032] Figure 2 It is a schematic flowchart of a cloud - native elastic GPU virtualization method for multi - resource mixed running of the present invention from the perspective of AI tasks;
[0033] Figure 3 It is a schematic diagram showing fine - grained segmentation of a cloud - native elastic GPU virtualization method for multi - resource mixed running of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0034] The present invention will now be described in further detail with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of them. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention. These drawings are all simplified schematic diagrams, only illustrating the basic structure of the present invention in a schematic manner, and therefore only showing the components related to the present invention.
[0035] To facilitate the understanding of the present invention, the present invention will be described more comprehensively below with reference to the relevant drawings. Several embodiments of the present invention are given in the drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. On the contrary, the purpose of providing these embodiments is to make the disclosure of the present invention more thorough and comprehensive.
[0036] All features disclosed in this specification, or all steps in the disclosed methods or processes, except for mutually exclusive features and / or steps, can be combined in any way.
[0037] Any feature disclosed in this specification (including any additional claims, abstract, and drawings), unless specifically stated, can be replaced by other equivalent or similar-purpose alternative features. That is, unless specifically stated, each feature is only an example in a series of equivalent or similar features.
[0038] In the present invention, unless otherwise clearly defined and limited, terms such as "installation", "connection", "connection", "fixation", etc. should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or integrated; it can be a mechanical connection, a direct connection, or an indirect connection through an intermediate medium. It can be the communication inside at least two elements or the interaction relationship between at least two elements, unless otherwise clearly defined. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0039] The following will be combined with Figures 1-3 to describe the present invention in detail. An embodiment provided by the present invention: A cloud-native elastic GPU virtualization method for multi-resource mixed running, including the following steps:
[0040] The first step: GPU computing power isolation and video memory isolation
[0041] First is the computing power isolation. Using the time-sharing mechanism, the virtualization management program slices the GPU computing resources according to time slices, and slices the usage rights of the GPU card into a series of time slices for sharing by the Pods on the k8sNode. If there is no GPU computing requirement for the current Pod within the time slice, this time slice will not be allocated to other Pods for use;
[0042] The hypervisor captures the CUDA program's submission events to the GPU through the kernel driver module. When initializing the kernel driver module, a hook function is registered to intercept the CUDA task submission operation. When the CUDA program calls the driver API to submit a task, the hook function is triggered to capture relevant events. The hook function records information such as the task type, parameters, and virtual machine ID. This mechanism allows the hypervisor to monitor the running status of GPU tasks in real-time and intervene in tasks that exceed the allocated specifications. If the running time of the CUDA program in the current time slice exceeds the Pod specification (for example, 50% of the GPU computing time is allocated, but the actual usage exceeds 50%), the hypervisor will set it to a blocked state and deprive it of GPU usage rights through the task scheduling subsystem of the Linux kernel to ensure that other Pods can fairly use GPU resources.
[0043] It should be noted that in order to reduce the overhead caused by time slice division, this solution provides a fine-grained time slice division scheme that supports 1% granularity GPU division. Similar to the GPU scheduling runlist, the GPU hypervisor generates a scheduling queue for all AI tasks on the Node. Each AI task can appear in the scheduling queue once or multiple times, and the minimum division can reach 1% of the GPU computing power. The fine-grained division scheme is determined according to user requirements. Refer to Figure 2 , and the scheduling process from the perspective of AI tasks is as follows:
[0044] ·init: The AI task starts (hijacks gpuregister through kernel driver ioctl) and enters the init state. The sGPU virtualization service routine allocates the corresponding time slice size for each AI task according to factors such as container specifications, priorities, and fine-grained scheduling algorithms, which is called the initial runnable time slice;
[0045] ·candidiate: At the beginning of each scheduling cycle, all AI tasks enter the candidate state;
[0046] ·active: The AI task enters the active state after using the GPU; the sGPU virtualization service routine regularly checks whether the time slice of the task in the active state has expired and updates the remaining runnable time slice;
[0047] ·throttled: When the remaining runnable time slice is less than or equal to 0, it means the time slice has expired, and the task enters the throttled state; it waits for the next scheduling cycle to arrive to regain schedulable time slices;
[0048] ·refill: After a complete scheduling cycle ends, the runnable time slice is restored to the initial runnable time slice;
[0049] deinit: When the AI task ends (hijacking gpu un-register through kernel driver ioctl) and enters the deinit state, the AI task is no longer managed by the sGPU from this point on.
[0050] Reference Figure 3 , the sGPU provides a configuration interface. Modifying the gpu.bandwidth file under the Pod directory can adjust the computing power specifications of the container.
[0051] Secondly, for video memory isolation, in the CUDA ecosystem, there are the following two ways of video memory allocation, but these two ways cannot be converted to each other;
[0052] · Non-uvm: cudaMalloc series of functions;
[0053] · uvm: cudaMallocManaged function;
[0054] The uvm method supports oversubscription of video memory. It can use system memory as the staging buffer to transfer data between system memory and video memory. However, the uvm allocation method is implemented through pagefault. When the kernel running on the GPU accesses data that is not in the video memory, it will trigger a pte exception event of GMMU. The GPU driver intervenes to transfer the data from system memory to video memory, and then notifies the hardware to re-execute that command. Therefore, the uvm allocation method has performance issues and is generally not used in production environments;
[0055] For the non-uvm method, video memory allocation is reserved. Excessive use of video memory by one task will cause other tasks to fail to run;
[0056] In this embodiment, the kernel driver module aggregates the usage of video memory in the Pod. When the usage of video memory exceeds the Pod specification, it directly returns a failure in video memory allocation:
[0057] Monitor the usage of video memory in the Pod through the kernel driver module to ensure that each Pod does not exceed its allocated video memory specification, aggregate the video memory usage in the Pod, and when the usage of video memory exceeds the Pod specification, directly return a failure in video memory allocation (similar to the behavior of memory allocation failure);
[0058] When initializing the kernel driver module, register a hook function to intercept video memory allocation and release operations. For example:
[0059] Intercept API calls such as cudaMalloc and cudaFree.
[0060] Record the allocated video memory amount, allocation time, and the PodID to which it belongs.
[0061] Monitor the video memory usage status
[0062] The kernel driver module can directly access the memory management unit of the GPU (such as UVM of NVIDIA or HSA of AMD) to obtain the real-time usage of the video memory.
[0063] For virtualized environments, the interfaces provided by the hypervisor (such as KVM or Xen) can be combined to obtain the video memory usage of each virtual machine (corresponding to a Pod).
[0064] Associate the Pod ID
[0065] In a Kubernetes environment, each Pod has a unique identifier (UID). The kernel driver module can obtain the Pod ID to which the current process belongs through the container runtime (such as Docker or containerd).
[0066] Associate the video memory usage information with the Pod ID for subsequent analysis.
[0067] Report the monitoring data
[0068] The kernel driver module can expose the captured video memory usage information to the user space through the / proc file system or the netlink interface.
[0069] Monitoring tools in the user space (such as Prometheus or Grafana) can regularly read this data and visualize it.
[0070] Video memory, like memory, is an incompressible resource and does not support elasticity even if there is enough free space on the GPU (for example, video memory requests exceeding the Pod specification will still be rejected);
[0071] Strictly limit the video memory usage of each Pod to avoid the failure of other Pods due to excessive video memory usage by a certain Pod, ensure that multiple Pods can fairly share the GPU video memory resources, and implement based on the existing kernel driver module without complex modification of the underlying hardware.
[0072] Step 2: The hypervisor provides elastic computing power scheduling capabilities and can use the remaining computing power beyond the committed virtual GPUs;
[0073] In the private cloud deployment scenario, the customer has the right to use the entire GPU card, but the actual utilization rate is very low. This embodiment provides the ability to elastically schedule the idle computing power of the GPU on the Node, and users can dynamically adjust and select the designed scheduling strategy according to specific scenarios.
[0074] The sGPU provides a kernel-level scheduling policy configuration interface gpu.sched_strategy, and the default value is 0, indicating fixed computing power. The designed scheduling strategies include the following:
[0075] Fixed computing power mode: echo "GPU_ID 0" > gpu.sched_strategy
[0076] The meaning of fixed computing power is that Pods on the Node use computing power according to the specifications, and the hypervisor strictly provides computing power according to the Pod specification declaration. Even if the current Pod has no GPU computing tasks, this time slice will not be allocated to other Pods for use;
[0077] Fixed computing power has the best effect on GPU computing power isolation, strictly restricting the use time of the GPU by AI tasks. The disadvantage is that the elastic ability of fixed computing power is the weakest, and it cannot fully utilize the idle computing resources of the GPU in the scenario where there is remaining computing power on the Node.
[0078] Dynamic computing power mode: echo "GPU_ID 1" > gpu.sched_strategy
[0079] The meaning of dynamic computing power is that all Pods on the Node evenly divide the idle computing power. If the total computing power specifications of all Pods in the current scheduling period are less than 100%, or some Pods have no GPU computing tasks during the scheduling time slice, the hypervisor considers that there is idle computing power in the current scheduling period, and it will evenly distribute the remaining time slice to each Pod for use;
[0080] Dynamic computing power is applicable to the scenario where there is remaining computing power on the Node. All Pods on the Node evenly distribute the idle computing resources of the GPU. Under dynamic computing power, the time for the container to use the GPU depends on two factors: idle computing power and the number of Pods with GPU computing tasks.
[0081] Minimum computing power mode: echo "GPU_ID 2" > gpu.sched_strategy
[0082] The meaning of minimum computing power is that all Pods on the Node will obtain at least the computing power resources declared by the Pod specification. If the total computing power specifications of all Pods in the current scheduling period are less than 100%, or some Pods have no GPU computing tasks during the scheduling time slice, the hypervisor considers that there is idle computing power in the current scheduling period, and it will randomly distribute the remaining time slice to the Pods for use;
[0083] Minimum computing power is applicable to the scenario where there is remaining computing power on the Node. The idle computing power of the GPU can be freely distributed. Under minimum computing power, the time for the container to use the GPU is uncertain, but at least it is the computing power specification applied for by the GPU.
[0084] Priority scheduling computing power mode: echo "GPU_ID3" > gpu.sched_strategy
[0085] The meaning of priority scheduling computing power is that oversubscription of computing power is provided on the Node. When the total computing power specifications of all Pods in the current scheduling cycle are greater than 100%, high-priority tasks can preempt the time slices of low-priority tasks. This solution provides an additional interface in the cgroup directory to configure the Pod priority;
[0086] Priority scheduling computing power is applicable to scenarios where there is remaining computing power on the Node. The idle computing power of the GPU is shared according to the priority order defined by the Pod. High-priority Pods use the idle computing power of the GPU first. Priority scheduling computing power is mainly applied in scenarios where AI tasks of multiple resource types run mixed.
[0087] The above four elastic scheduling strategies are all based on the configuration files of the linux kernel sys file system. The adjustment of the elastic computing power strategy takes effect in the next scheduling cycle, and the effective cycle is in microseconds.
[0088] The third step: · Container SLA guarantee, sense the health of business containers and the GPU load situation, feedback to the apiserver through kubelet to adjust the cluster scheduling strategy, and perform dynamic node scheduling on Pods;
[0089] Container health perception:
[0090] Use the built-in probes of Kubernetes (such as LivenessProbe, ReadinessProbe) to detect the health status of containers, and collect the usage of resources such as CPU, memory, and network of containers.
[0091] GPU load perception:
[0092] Use GPU performance monitoring tools (such as NVIDIA DCGM, AMD ROCm) to obtain indicators such as the usage rate of the GPU, video memory occupancy, and temperature, and bind these indicators to the containers to form a GPU load view for each Pod.
[0093] Summarize the health status and GPU load information of all Pods at the Node level (through Kubelet or extended components), and convert this information into structured indicators (such as JSON format) for subsequent processing.
[0094] SLA evaluation:
[0095] Evaluate whether each Pod meets its SLA requirements according to predefined SLA rules (such as minimum GPU computing power requirements, maximum latency tolerance, etc.). If a certain Pod does not meet the SLA, mark it as a state that requires intervention.
[0096] Kubelet reports:
[0097] Kubelet reports the aggregated Pod health and GPU load information to the APIServer, and this information can be passed through custom CustomResourceDefinition (CRD) or event mechanism.
[0098] Scheduling policy adjustment:
[0099] The APIServer dynamically adjusts the scheduling policy according to the received information;
[0100] Scheduling based on health status: If the health status of a certain Pod is poor, try to reschedule it to other nodes.
[0101] Scheduling based on GPU load: If the GPU load on a certain Node is too high, preferentially schedule new Pods to idle nodes.
[0102] Priority adjustment: For low-priority tasks, they can be actively evicted to release resources.
[0103] This embodiment is a GPU virtualization cloud-native solution based on the k8s scheduling framework. It provides a GPU virtualization method at the single-machine level. Through kubelet reporting the running status of Pods and the load on Nodes, the ApiServer of k8s dynamically adjusts the elastic policy configuration on Nodes by aggregating the workload of the cluster and the historical resource profiles. The update of the entire link can take effect in a short time.
[0104] To sum up, the solution of the present invention takes cost into account and combines well with the cloud-native ecosystem, making the GPU computing power optimal in terms of supporting serviceization, flexible computing power scheduling, and effective utilization rate. This solution abstracts the execution environment of AI tasks and provides virtual GPU products with specific computing power and video memory specifications. Each service running on the virtual GPU card thinks it has exclusive access to the entire card resources. This solution also provides isolation technologies for computing power and video memory to ensure the usage experience and fault isolation ability of AI tasks.
[0105] The resource isolation technology of GPU virtualization provides the infrastructure for sharing GPU resources in a multi-tenant scenario. In the private cloud deployment scenario, customers have the right to use the entire GPU card, but the actual utilization rate is very low. Customers hope to ensure resource isolation between Pods while fully utilizing the computing power resources of the GPU card.
[0106] Based on achieving the isolation of computing power and video memory and decoupling the user from the physical card, this solution provides the ability of multi-task elastic computing power mixed running to further optimize the scenario of low overall card utilization; the computing power other than the allocated virtual GPU will be allocated to all Pods on the Node according to certain rules.
[0107] During the window period with abundant resources, the ability of elastic computing power collaboration is added to improve the actual utilization rate of the GPU card. It makes it possible for multi-resource type tasks to run mixed, greatly improving the resource efficiency and reducing the usage cost of the GPU card through technological dividends.
[0108] The above details an embodiment corresponding to a cloud-native elastic GPU virtualization method for multi-resource mixed running. Based on this, the present invention also discloses a cloud-native elastic GPU virtualization system and a storage medium corresponding to the above method.
[0109] A cloud-native elastic GPU virtualization system for multi-resource mixed running includes:
[0110] A memory for storing computer programs;
[0111] A processor for executing the computer programs, and when the computer programs are executed by the processor, they can implement the relevant steps in a cloud-native elastic GPU virtualization method for multi-resource mixed running disclosed in any of the foregoing embodiments.
[0112] Among them, the processor may include one or more processing cores, such as a core processor, a core processor, etc. The processor can be implemented in at least one hardware form of a digital signal processor DSP (Digital Signal Processing), a field programmable gate array FPGA (Field-Programmable Gate Array), and a programmable logic array PLA (Programmable Logic Array). The processor may also include a main processor and a coprocessor. The main processor is a processor used to process data in the wake state, also known as a central processing unit CPU (Central Processing Unit); the coprocessor is a low-power processor used to process data in the standby state.
[0113] In some embodiments, the processor may be integrated with a graphics processing unit GPU (Graphics Processing Unit), and the GPU is responsible for the rendering and drawing of the content to be displayed on the display screen. In some embodiments, the processor may also include an artificial intelligence AI (Artificial Intelligence) processor, and the AI processor is used to process computational operations related to machine learning.
[0114] The memory may include one or more readable storage media, which may be non-transitory. The memory may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In this embodiment, the memory is at least used to store the following computer program, wherein, after the computer program is loaded and executed by the processor, it can implement the relevant steps in a multi-resource mixed-running cloud-native elastic GPU virtualization method disclosed in any of the foregoing embodiments. In addition, the resources stored in the memory may also include an operating system and data, etc., and the storage method may be transient storage or permanent storage. Among them, the operating system may be Windows. The data may include, but is not limited to, the data involved in the above method.
[0115] In addition, in each embodiment of the present invention, each functional module may be integrated in a processing module, may exist physically separately for each module, or two or more modules may be integrated in one module. The above integrated module may be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may be stored in a readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of this technical solution, may be embodied in the form of a software product, and this computer software product is stored in a storage medium and executes all or part of the steps of the methods described in each embodiment of the present invention.
[0116] Therefore, an embodiment of the present invention further provides a readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of a multi-resource mixed-running cloud-native elastic GPU virtualization method.
[0117] The readable storage medium may include: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memory ROM (Read-Only Memory), random access memory RAM (RandomAccessMemory), magnetic disks or optical discs.
[0118] The computer program included in the readable storage medium provided in this embodiment can implement the steps of a multi-resource mixed-running cloud-native elastic GPU virtualization method as described above when executed by the processor, and the effect is the same as above.
[0119] The above has introduced in detail a cloud-native elastic GPU virtualization method, system and storage medium for multi-resource mixed running provided by the present invention. Each embodiment in the specification is described in a progressive manner, and the key point of each embodiment is the difference from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the devices, equipment and readable storage media disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and modifications can be made to the present invention, and these improvements and modifications also fall within the protection scope of the claims of the present invention.
[0120] As described above, it is only the specific implementation manner of the invention, but the protection scope of the invention is not limited thereto. Any change or replacement that can be thought of without creative work should be covered within the protection scope of the invention. Therefore, the protection scope of the invention should be subject to the protection scope defined by the claims.
Claims
1. A cloud-native elastic GPU virtualization method for multi-resource mixed running, characterized in that: It includes fine-grained time slicing of GPU computing resources, where each sliced time slice is allocated to the corresponding Pod for use, and the kernel-mode driver is used to monitor the running status of GPU tasks in real time. Intervention is performed on Pods whose running time and video memory usage exceed the allocated specifications. The GPU computing power scheduling strategy designed is dynamically adjusted according to the usage scenario. At the same time, the container SLA guarantee mechanism is used to sense the health of business containers and the GPU load situation, and dynamic node scheduling is performed on Pods.
2. The cloud-native elastic GPU virtualization method for multi-resource mixed running according to claim 1, wherein: The fine-grained time slicing supports 1% granularity GPU slicing. The usage right of the GPU card is sliced into a series of time slices and shared by Pods on the k8s Node. If there is no GPU computing requirement for the current Pod within a time slice, this time slice will not be allocated to other Pods for use.
3. A cloud-native elastic GPU virtualization method for multi-resource mixed running according to claim 2, characterized in that: The method for intervening in Pods whose running time exceeds the allocated specifications includes that the hypervisor marks the CUDA program running within the current time slice in a blocked state for Pods whose running time exceeds their allocated specifications, and then the GPU usage right of the Pod is deprived through the task scheduling subsystem of the Linux kernel.
4. The method for cloud-native elastic GPU virtualization with multi-resource mixed running according to claim 3, wherein: The method for intervening in Pods whose video memory usage exceeds the allocated specifications includes monitoring and summarizing the video memory usage in the Pod through the kernel driver module. When the video memory usage exceeds the Pod specifications, a video memory allocation failure is directly returned.
5. A cloud-native elastic GPU virtualization method for multi-resource mixed running according to claim 4, characterized in that: The scheduling strategy includes: Fixed computing power mode, where Pods on the Node use computing power according to the specifications, and the hypervisor strictly provides computing power according to the Pod specifications declared. Even if the current Pod has no GPU computing tasks, this time slice will not be allocated to other Pods for use; Dynamic computing power mode, where all Pods on the Node equally divide the idle computing power. If the total computing power specifications of all Pods within the current scheduling cycle are less than 100%, or some Pods have no GPU computing tasks within the scheduling time slice, the hypervisor considers that there is idle computing power in the current scheduling cycle, and it will evenly allocate the remaining time slices to each Pod for use; Minimum computing power mode, where all Pods on the Node will obtain at least the computing power resources declared by the Pod specifications. If the total computing power specifications of all Pods within the current scheduling cycle are less than 100%, or some Pods have no GPU computing tasks within the scheduling time slice, the hypervisor considers that there is idle computing power in the current scheduling cycle, and it will randomly allocate the remaining time slices to the Pods for use; Priority scheduling computing power mode, where oversubscription of computing power is provided on the Node. When the total computing power specifications of all Pods within the current scheduling cycle are greater than 100%, high-priority tasks can preempt the time slices of low-priority tasks. This solution provides an additional interface in the cgroup directory to configure the Pod priority.
6. A cloud-native elastic GPU virtualization method for multi-resource mixed running according to claim 5, characterized in that: The container SLA guarantee mechanism will sense the health of business containers and the GPU load situation and feedback it to the apiserver through the kubelet to adjust the cluster scheduling strategy.
7. A cloud-native elastic GPU virtualization method for multi-resource mixed running according to claim 6, characterized in that: The Pod dynamic node scheduling includes: Scheduling based on health status. If the health status of a Pod is poor, it attempts to reschedule it to other nodes; GPU load-based scheduling, where the GPU load on one Node is too high, and new Pods are preferentially scheduled to idle nodes; Priority-based adjustment: For low-priority tasks, they can be actively evicted to release resources.
8. A cloud-native elastic GPU virtualization method for multi-resource mixed running according to claim 1, characterized in that: Monitoring the running status of GPU tasks includes that the hypervisor intercepts the event of a CUDA program submitting a task to the GPU through the kernel driver module. When initializing the kernel driver module, a hook function is registered to intercept the CUDA task submission operation. When the CUDA program calls the driver API to submit a task, the hook function will be triggered to capture relevant events, and the hook function records information such as the type of the task, parameters, and virtual machine ID.
9. A cloud-native elastic GPU virtualization system for multi-resource mixed running, characterized in that: Including: A memory for storing computer programs; A processor for executing the computer program, and when the computer program is executed by the processor, it implements the steps of a cloud-native elastic GPU virtualization method for multi-resource mixed running as described in any one of claims 1-8.
10. A readable storage medium, on which a computer program is stored, characterized in that: When the computer program is executed by the processor, it implements the steps of a cloud-native elastic GPU virtualization method for multi-resource mixed running as described in any one of claims 1-8.