GPU task execution method, device and equipment based on Kubernetes and medium

By dividing virtual GPU units in the Kubernetes cluster and using the long short-term memory model to predict task loads, combined with intelligent scheduling and video memory optimization technology, the problem of inefficient GPU resource management in a multi-tenant environment is solved, efficient sharing and refined allocation are achieved, and resource utilization and task execution efficiency are improved.

CN120723458APending Publication Date: 2025-09-30INSPUR ENTERPRISE CLOUD TECHNOLOGY (SHANDONG) CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510875243.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2025-09-30

AI Technical Summary

Technical Problem

Kubernetes cannot effectively manage and schedule GPU resources in a multi-tenant environment, resulting in inefficient resource utilization and a lack of a refined isolation mechanism for computing cores and video memory resources.

Method used

By dividing virtual GPU units in the Kubernetes cluster, using the long short-term memory model to predict the task load curve, and combining intelligent scheduling and memory optimization technology, resources are dynamically allocated to meet the needs of different tasks.

Benefits of technology

It achieves efficient sharing and refined allocation of GPU resources in a multi-tenant environment, improves resource utilization and task execution efficiency, and ensures task isolation and fairness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120723458A_ABST
    Figure CN120723458A_ABST
Patent Text Reader

Abstract

The invention discloses a GPU task execution method and device based on Kubernetes, equipment and a medium, and relates to the technical field of GPU virtualization, and the method comprises the steps: determining GPU resource demand information of a target GPU task based on a preset scheduling expander in a Kubernetes cluster, and evaluating the resource load of each GPU node in the Kubernetes cluster; predicting a task load curve of the target GPU task by using the target load prediction model, and generating a target time slice strategy instruction of the target GPU task based on the task load curve; determining a target GPU node based on the task load curve, the GPU resource demand information and the resource load of each GPU node; and determining a target virtual GPU unit in the target GPU node, and executing the target GPU task based on the target time slice strategy instruction by using the target virtual GPU unit. According to the invention, efficient sharing and refined allocation of GPU resources can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of GPU virtualization, and in particular to a Kubernetes-based GPU task execution method, apparatus, device, and medium. Background Art

[0002] With the rapid development of cloud computing and containerization technologies, Kubernetes (K8s) has become a core platform for large-scale cluster resource management in modern data centers. As a containerized management tool, Kubernetes' flexibility and scalability provide powerful support for resource scheduling, automated management, and elastic scaling. However, despite Kubernetes's excellent performance in managing resources such as CPU and memory, it still faces numerous challenges in managing and scheduling GPU (Graphics Processing Unit) resources. Currently, Kubernetes primarily allocates GPU resources on a per-GPU basis. This coarse-grained allocation method cannot meet the resource requirements of different tasks in a multi-tenant environment. Furthermore, in a multi-tenant environment, multiple tasks may share the same physical GPU resources, and most existing GPU virtualization technologies fail to effectively guarantee the independence of compute cores and memory resources, lacking refined isolation mechanisms.

[0003] In summary, how to achieve efficient sharing and refined allocation of GPU resources in a multi-tenant environment is a technical problem that needs to be solved urgently. Summary of the Invention

[0004] In view of this, the purpose of the present invention is to provide a Kubernetes-based GPU task execution method, apparatus, device, and medium that can achieve efficient sharing and refined allocation of GPU resources in a multi-tenant environment. The specific solution is as follows:

[0005] In a first aspect, the present application provides a GPU task execution method based on Kubernetes, comprising:

[0006] Obtain a target GPU task, determine the GPU resource requirement information of the target GPU task based on a preset scheduling extender in the Kubernetes cluster, and evaluate the resource load of each GPU node in the Kubernetes cluster;

[0007] A target load prediction model is used to predict a task load curve of the target GPU task, and a target time slice policy instruction corresponding to the target GPU task is generated based on the task load curve; the target load prediction model is a model obtained by training based on a long short-term memory model;

[0008] Determine a target GPU node in the Kubernetes cluster that meets the execution condition corresponding to the target GPU task based on the task load curve, the GPU resource requirement information, and the resource load of each GPU node;

[0009] A target virtual GPU unit in the target GPU node is determined, and the target virtual GPU unit is used to execute the target GPU task based on the target time slice policy instruction.

[0010] Optionally, the Kubernetes-based GPU task execution method further includes:

[0011] Deploy a preset device plug-in on each GPU node in the Kubernetes cluster, divide each GPU node into a plurality of virtual GPU units based on a preset business scenario using the preset device plug-in, and register each virtual GPU unit with the Kubernetes cluster using a target remote procedure call framework;

[0012] The virtual GPU resource attributes and working status of each virtual GPU unit in each GPU node are defined in the Kubernetes cluster through a preset resource management mechanism to obtain a corresponding target GPU resource management table; the virtual GPU resource attributes include computing cores and video memory, and the working status includes idle and occupied.

[0013] Optionally, evaluating the resource load of each GPU node in the Kubernetes cluster based on a preset scheduling extender in the Kubernetes cluster includes:

[0014] Determine the number of first virtual GPU units in idle working state in each GPU node based on the preset scheduling extender and the target GPU resource management table in the Kubernetes cluster, and determine the target number of idle computing cores and first idle video memory corresponding to the first virtual GPU unit;

[0015] Determine the second idle video memory corresponding to each GPU node based on an allocation table in a preset video memory optimizer, and determine a corresponding target idle video memory based on the first idle video memory and the second idle video memory;

[0016] The resource load of each GPU node in the Kubernetes cluster is evaluated based on the target number of idle computing cores, the target idle video memory, and the number of the first virtual GPU units.

[0017] Optionally, before predicting the task load curve of the target GPU task using the target load prediction model, the method further includes:

[0018] Determine a task type corresponding to the target GPU task, obtain several historical GPU tasks of the same task type, and obtain historical task load data corresponding to each of the historical GPU tasks; the task types include inference tasks and training tasks;

[0019] Using the long short-term memory model to learn and analyze the historical task load data corresponding to each of the historical GPU tasks to obtain the target load prediction model;

[0020] Accordingly, the method of predicting the task load curve of the target GPU task using the target load prediction model and generating the target time slice policy instruction corresponding to the target GPU task based on the task load curve includes:

[0021] Inputting the GPU resource requirement information corresponding to the target GPU task into the target load prediction model, and outputting the task load curve corresponding to the target GPU task using the target load prediction model;

[0022] The target time slice policy instruction corresponding to the target GPU task is generated based on the task type corresponding to the target GPU task and the task load curve.

[0023] Optionally, the Kubernetes-based GPU task execution method further includes:

[0024] If it is determined based on the task load curve, the GPU resource requirement information and the resource load of each GPU node that there is no target GPU node in the Kubernetes cluster that meets the execution conditions corresponding to the target GPU task, the target priority corresponding to the target GPU task is determined based on the task type corresponding to the target GPU task, and the target GPU task is added to the target to-be-executed GPU task queue based on the target priority, so that each to-be-executed GPU task is executed based on the priority corresponding to each to-be-executed GPU task in the target to-be-executed GPU task queue.

[0025] Optionally, before using the target virtual GPU unit to execute the target GPU task based on the target time slice policy instruction, the method further includes:

[0026] Deploy a multi-process service daemon on each GPU node in the Kubernetes cluster;

[0027] Creating a computing context corresponding to the target GPU task by using the multi-process service daemon on the target GPU node, and obtaining a target time slice policy instruction corresponding to the target GPU task by using the multi-process service daemon;

[0028] Accordingly, the using the target virtual GPU unit to execute the target GPU task based on the target time slice policy instruction includes:

[0029] Bind the target virtual GPU unit to the target GPU task, and determine the target video memory constraint condition corresponding to the target GPU task;

[0030] The target GPU task is executed using the target virtual GPU unit based on the target video memory constraint, the computing context corresponding to the target GPU task, and the target time slice policy instruction.

[0031] Optionally, the process of executing the target GPU task using the target virtual GPU unit based on the target video memory constraint, the computing context corresponding to the target GPU task, and the target time slice policy instruction further includes:

[0032] monitoring a fragmentation rate of the video memory using an allocation table in a preset video memory optimizer; determining a target video memory compression rate if the fragmentation rate exceeds a fragmentation rate threshold corresponding to the target video memory constraint; compressing the video memory of the target virtual GPU unit based on the target video memory compression rate using a target compression engine; and updating the allocation table based on the compressed video memory;

[0033] An actual execution time slice of the target GPU task is monitored, and the target load prediction model is updated based on the target GPU task and the actual execution time slice.

[0034] In a second aspect, the present application provides a Kubernetes-based GPU task execution device, comprising:

[0035] A resource assessment module is used to obtain a target GPU task, determine the GPU resource requirement information of the target GPU task based on a preset scheduling extender in the Kubernetes cluster, and assess the resource load of each GPU node in the Kubernetes cluster;

[0036] a time slice policy instruction generation module, configured to predict a task load curve of the target GPU task using a target load prediction model, and generate a target time slice policy instruction corresponding to the target GPU task based on the task load curve; the target load prediction model is a model trained based on a long short-term memory model;

[0037] A target GPU node determination module is configured to determine a target GPU node in the Kubernetes cluster that meets the execution conditions corresponding to the target GPU task based on the task load curve, the GPU resource requirement information, and the resource load of each GPU node;

[0038] The target GPU task execution module is configured to determine a target virtual GPU unit in the target GPU node, and utilize the target virtual GPU unit to execute the target GPU task based on the target time slice policy instruction.

[0039] In a third aspect, the present application provides an electronic device, comprising:

[0040] Memory, used to store computer programs;

[0041] A processor is configured to execute the computer program to implement the aforementioned Kubernetes-based GPU task execution method.

[0042] In a fourth aspect, the present application provides a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, the aforementioned Kubernetes-based GPU task execution method is implemented.

[0043] In this application, a target GPU task is first obtained, and the GPU resource requirement information of the target GPU task is determined based on the preset scheduling extender in the Kubernetes cluster, and the resource load of each GPU node in the Kubernetes cluster is evaluated; then, a target load prediction model is used to predict the task load curve of the target GPU task, and a target time slice policy instruction corresponding to the target GPU task is generated based on the task load curve; the target load prediction model is a model obtained by training based on a long short-term memory model; then, based on the task load curve, the GPU resource requirement information and the resource load of each GPU node, the target GPU node in the Kubernetes cluster that meets the execution conditions corresponding to the target GPU task is determined; finally, a target virtual GPU unit in the target GPU node is determined, and the target virtual GPU unit is used to execute the target GPU task based on the target time slice policy instruction. As can be seen from the above, in this application, the preset scheduling extender in the Kubernetes cluster is first used to determine the GPU resource requirement information of the target GPU task and evaluate the resource load of each GPU node in the Kubernetes cluster; then, the target load prediction model obtained by training the long short-term memory model is used to predict the task load curve of the target GPU task, and the target time slice policy instruction of the target GPU task is generated according to the task load curve; then, the target GPU node in the Kubernetes cluster that meets the execution conditions of the target GPU task is determined according to the task load curve, GPU resource requirement information, and the resource load of each GPU node; finally, the target virtual GPU unit in the target GPU node is determined, and the target virtual GPU unit is used to execute the target GPU task based on the target time slice policy instruction. In this way, the present application can refine the allocation of GPU resources in the Kubernetes cluster by dividing the physical GPU into virtual GPU units, and then using the virtual GPU units in the GPU nodes that meet the GPU task execution conditions to execute GPU tasks, thereby avoiding waste and improving resource utilization efficiency. In this way, the present application can achieve efficient sharing and refined allocation of GPU resources in a multi-tenant environment, has strong scalability, and meets the needs of GPU virtualization management and resource scheduling in large-scale cloud computing environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.

[0045] Figure 1 A Kubernetes-based GPU task execution system architecture diagram provided for this application;

[0046] Figure 2 A flowchart of a Kubernetes-based GPU task execution method provided in this application;

[0047] Figure 3 A specific virtual GPU unit division diagram provided in this application;

[0048] Figure 4 A specific workflow diagram of the execution optimization layer provided in this application;

[0049] Figure 5 A flowchart of a specific Kubernetes-based GPU task execution method provided in this application;

[0050] Figure 6 A schematic diagram of the structure of a GPU task execution device based on Kubernetes provided in this application;

[0051] Figure 7 This is a structural diagram of an electronic device provided in this application. DETAILED DESCRIPTION

[0052] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0053] Currently, Kubernetes schedules GPU resources primarily on a per-GPU basis. This coarse-grained allocation method cannot meet the resource requirements of different tasks in a multi-tenant environment. Furthermore, in a multi-tenant environment, multiple tasks may share the same physical GPU resources, and most existing GPU virtualization technologies fail to effectively guarantee the independence of computing cores and video memory resources, lacking refined isolation mechanisms. Therefore, this application provides a Kubernetes-based GPU task execution solution that enables efficient sharing and refined allocation of GPU resources in a multi-tenant environment.

[0054] The system framework used in this application's Kubernetes-based GPU task execution solution can be found in Figure 1As shown, specifically, this application includes a virtualization resource management layer, an intelligent scheduling layer, a dynamic time-sharing control layer, and an execution optimization layer. The virtualization resource management layer consists of a GPU device plug-in and GPU CRD (Custom Resource Definitions), which are used to implement virtualization and standardization of GPU resources. The intelligent scheduling layer includes a SchedulerExtender and a resource monitoring module, which are responsible for resource scheduling decisions. The dynamic time-sharing control layer is equipped with an LSTM (Long Short-Term Memory) predictor and a time slice allocator, which are used to predict the load of GPU tasks and dynamically allocate computing resources. The execution optimization layer includes an MPS service (Multi-Process Service) and a memory compression module, which are used to optimize task execution efficiency. The layers interact through data streams to form a complete closed-loop GPU resource management system. Among them, the GPU device plug-in is connected to the GPU CRD to register virtual resources, the GPU CRD is connected to the SchedulerExtender to provide resource information, the Scheduler Extender is connected to the resource monitoring module, LSTM predictor and MPS service respectively to realize scheduling decision, prediction analysis and task distribution, the LSTM predictor is connected to the time slice allocator to provide load prediction results, the time slice allocator is connected to the Scheduler Extender to return the allocation strategy, the MPS service is connected to the memory compression module to optimize execution efficiency, and the memory compression module is connected to the resource monitoring module to feedback optimization indicators.

[0055] See also Figure 2 As shown, an embodiment of the present invention discloses a GPU task execution method based on Kubernetes, which may include:

[0056] Step S11: Obtain a target GPU task, determine the GPU resource requirement information of the target GPU task based on a preset scheduling extender in the Kubernetes cluster, and evaluate the resource load of each GPU node in the Kubernetes cluster.

[0057] In this embodiment, see Figure 3As shown, to achieve multi-tenant GPU resource sharing, the virtualized resource management layer must first virtualize and manage GPU resources. Traditional GPU resource allocation methods typically allocate resources in a single manner, resulting in low resource utilization efficiency and failing to fully meet the needs of different tasks in a multi-tenant environment. To address this issue, this embodiment implements virtualized GPU resource management through a custom GPU device plugin and CRD. Specifically, this may include: first deploying a preset device plugin on each GPU node in the Kubernetes cluster, using the preset device plugin to divide each GPU node into several virtual GPU units based on preset business scenarios, and registering each virtual GPU unit with the Kubernetes cluster using a target remote procedure call framework; then, defining the virtual GPU resource attributes and operating status of each virtual GPU unit in each GPU node in the Kubernetes cluster using a preset resource management mechanism to obtain a corresponding target GPU resource management table; virtual GPU resource attributes include computing cores and video memory, and the operating status includes idle and occupied. Specifically, a custom GPU device plugin is installed and run on each GPU node in the Kubernetes cluster. This plugin allows users to partition a physical GPU into multiple virtual GPU units, each consisting of a portion of compute cores and video memory. The plugin communicates with Kubernetes via gRPC (Remote Procedure Call), registering each virtual GPU unit with the Kubernetes cluster so that the scheduler can identify and allocate these resources. Virtual GPU units are then defined using CRDs. For example, nvidia.com / gpu-core represents the number of compute cores in a virtual GPU unit, and nvidia.com / gpu-mem represents the memory size of the virtual GPU unit. This allows the Kubernetes scheduler to identify and schedule the resources of these virtual GPU units. The scheduler can then accurately match the required resources of a task with the corresponding virtual GPU unit based on its specific requirements, achieving more refined and efficient resource management. This virtualized resource management method not only effectively allocates GPU computing power and memory bandwidth to multiple tenants, but also ensures that each tenant receives appropriate resources based on the needs of its tasks and avoids resource contention between tasks.

[0058] It should be noted that to ensure efficient allocation of GPU resources in a multi-tenant environment, the intelligent scheduling layer optimizes GPU resource scheduling and management by extending the Kubernetes cluster's scheduler. This custom Kubernetes SchedulerExtender dynamically selects the most suitable GPU node for task scheduling based on the GPU resource requirements of the GPU task (such as the number of compute cores and video memory size) and the GPU resource load of each GPU node in the Kubernetes cluster, ensuring efficient and fair allocation of GPU resources in a multi-tenant environment.

[0059] In this embodiment, the above-mentioned evaluation of the resource load of each GPU node in the Kubernetes cluster based on the preset scheduler extender in the Kubernetes cluster may include: first, determining the number of first virtual GPU units in the idle working state in each GPU node based on the preset scheduler extender in the Kubernetes cluster and the target GPU resource management table, and determining the target number of idle computing cores and first idle video memory corresponding to the first virtual GPU unit; then, determining the second idle video memory corresponding to each GPU node based on the allocation table in the preset video memory optimizer, and determining the corresponding target idle video memory based on the first idle video memory and the second idle video memory; finally, evaluating the resource load of each GPU node in the Kubernetes cluster based on the target number of idle computing cores, the target idle video memory and the number of the first virtual GPU units. Specifically, the Scheduler Extender first performs a detailed evaluation of the resource requirements of the target GPU task, mainly including the number of computing cores, video memory size, etc. Then, the Scheduler Extender makes a decision based on the resource load of each GPU node in the Kubernetes cluster, especially the number of idle computing cores, video memory bandwidth and current load on each GPU node. The Scheduler Extender can dynamically select the most suitable GPU node to execute the target GPU task based on these resource data. Through this process, tasks can be scheduled to the GPU nodes with the most idle and suitable resources, thereby achieving load balancing. In addition, in order to respond to changes in GPU resources in real time, Scheduler Extender will continuously monitor the GPU resource usage of each GPU node. Scheduler Extender will regularly obtain the resource status of each GPU node, including the number of idle computing cores of each GPU node, video memory usage, and the resource load of the current GPU node. These resource data will help Scheduler Extender make real-time scheduling decisions to avoid GPU resource overload of certain GPU nodes, thereby affecting the execution efficiency of tasks.

[0060] It should be noted that Scheduler Extender also has the ability to dynamically adjust GPU resource allocation based on task load. When a task's computational requirements change, Scheduler Extender dynamically adjusts resource allocation. For example, for tasks with increased resource requirements, Scheduler Extender can expand the task's resource quota or adjust the task's time slice length to adapt to the current load. This dynamic adjustment mechanism ensures timely response and support for tasks as resource requirements change. As can be seen, by combining task type, priority, resource requirements, and node resource status, Scheduler Extender can precisely manage the allocation of each virtual GPU unit. This maximizes GPU resource utilization and avoids performance bottlenecks caused by resource contention. Furthermore, Scheduler Extender ensures resource isolation between multi-tenant tasks, preventing resource competition between different tenants and achieving fair resource allocation. As a result, the Scheduler Extender, through its pre-configured scheduler, can efficiently schedule GPU resources in a multi-tenant environment, ensuring the performance requirements of each tenant's task and reducing resource contention and the waste of idle resources, thereby improving task computational efficiency and GPU resource utilization.

[0061] Step S12: using a target load prediction model to predict a task load curve of the target GPU task, and generating a target time slice policy instruction corresponding to the target GPU task based on the task load curve; the target load prediction model is a model obtained by training based on a long short-term memory model.

[0062] In a multi-tenant environment, efficient allocation and management of GPU resources is a key factor in improving resource utilization. Different types of tasks have significantly different requirements for GPU resources. For example, training tasks typically require longer computation times and higher computational resources, while inference tasks have higher requirements for response time and relatively less computational effort. Therefore, the dynamic time-sharing control layer in this embodiment can utilize load prediction technology, combined with an adaptive time-slice allocation mechanism, to improve GPU resource utilization efficiency and meet the resource requirements of different tasks.

[0063] It should be noted that before using the target load prediction model to predict the task load curve of the target GPU task, the method may also include: first determining the task type corresponding to the target GPU task, and obtaining several historical GPU tasks of the same task type, and obtaining historical task load data corresponding to each of the historical GPU tasks; the task types include inference tasks and training tasks; and then using the long short-term memory model to learn and analyze the historical task load data corresponding to each of the historical GPU tasks to obtain the target load prediction model. Specifically, the LSTM model is used to learn and analyze the historical task load data of historical tasks of the same task type as the target GPU task to obtain a target load prediction model, and the target load prediction model is used to predict the GPU load changes of the task. The target load prediction model can accurately predict the computing requirements of the task, obtain the task load curve, and dynamically adjust the time slice allocation of the task based on the prediction results.

[0064] In this embodiment, using a target load prediction model to predict the target GPU task's task load curve and generating a target time slice policy instruction corresponding to the target GPU task based on the task load curve may include: first, inputting the GPU resource requirement information corresponding to the target GPU task into the target load prediction model, and using the target load prediction model to output the task load curve corresponding to the target GPU task; and finally, generating the target time slice policy instruction corresponding to the target GPU task based on the task type and task load curve corresponding to the target GPU task. Specifically, for computationally intensive tasks requiring longer GPU computation time, the system allocates longer time slices; whereas for tasks with higher response time requirements and lower computational load, the system allocates shorter time slices. Furthermore, different tasks can be executed alternately based on their assigned time slices. As can be seen, the LSTM model can predict future task load requirements in real time by analyzing historical task data. When the task load is high, the system automatically increases the time slice length based on the predicted results to ensure successful task completion. When the task load is low, the system correspondingly reduces the allocated time slice to avoid wasting GPU resources. In this way, the system can adaptively adjust resource allocation according to changes in task load and maximize the utilization of GPU resources.

[0065] In this embodiment, by implementing the adaptive time-slice allocation feature, resources can be dynamically adjusted based on the needs of different task types, ensuring efficient use of GPU resources during task execution while reducing unnecessary resource waste. This mechanism not only improves GPU resource utilization efficiency but also enhances fairness and isolation between tasks in a multi-tenant environment.

[0066] Step S13: Determine a target GPU node in the Kubernetes cluster that meets the execution condition corresponding to the target GPU task based on the task load curve, the GPU resource requirement information, and the resource load of each GPU node.

[0067] In this embodiment, the intelligent scheduling layer evaluates the resource load of each GPU node in the Kubernetes cluster based on the task load curve and GPU resource requirement information corresponding to the target GPU task, and determines whether there is a target GPU node in the Kubernetes cluster that meets the execution conditions corresponding to the target GPU task. It should be noted that if, based on the task load curve, GPU resource requirement information, and resource load of each GPU node, it is determined that no target GPU node in the Kubernetes cluster meets the execution conditions corresponding to the target GPU task, the target priority of the target GPU task is determined based on the task type corresponding to the target GPU task. Based on the target priority, the target GPU task is added to the target pending GPU task queue, so that each pending GPU task in the target pending GPU task queue is executed based on its corresponding priority. Specifically, in a multi-tenant environment, task priority is a key factor in determining resource allocation for tasks in the target pending GPU task queue. The Scheduler Extender dynamically adjusts the scheduling order of tasks based on their priority. When multiple tasks request resources, the system prioritizes high-priority tasks to ensure the timely execution of critical tasks. This priority scheduling strategy ensures that critical business tasks are not blocked by low-priority tasks, ensuring high system availability and timely task completion.

[0068] Step S14: Determine a target virtual GPU unit in the target GPU node, and use the target virtual GPU unit to execute the target GPU task based on the target time slice policy instruction.

[0069] To efficiently share GPU resources and ensure task isolation and security in a multi-tenant environment, the execution optimization layer uses CUDA MPS technology to optimize GPU resource allocation and management. It's important to note that in traditional GPU resource allocation, each process exclusively uses a GPU context, which results in frequent context switching and impacts task execution efficiency. MPS technology allows multiple CUDA applications to share the computing resources of the same physical GPU, reducing context switching and improving overall resource utilization. In addition to task context isolation, the MPS service also incorporates video memory compression technology. Video memory is a limited and expensive GPU resource. Especially in a multi-tenant environment, when multiple tasks use GPU resources simultaneously, video memory fragmentation can occur, leading to low memory utilization. To improve memory efficiency, the MPS service uses a dynamic memory compression mechanism to compress idle video memory, reducing memory fragmentation and optimizing memory management.

[0070] In this embodiment, before executing the target GPU task based on the target time slice policy instruction using the target virtual GPU unit, the method may further include: first deploying a multi-process service daemon on each GPU node in the Kubernetes cluster; then creating a computing context corresponding to the target GPU task using the multi-process service daemon on the target GPU node, and obtaining the target time slice policy instruction corresponding to the target GPU task using the multi-process service daemon. Accordingly, executing the target GPU task based on the target time slice policy instruction using the target virtual GPU unit may include: first binding the target virtual GPU unit to the target GPU task and determining the target memory constraint corresponding to the target GPU task; then executing the target GPU task based on the target memory constraint, the computing context corresponding to the target GPU task, and the target time slice policy instruction using the target virtual GPU unit. Specifically, an MPS daemon is deployed on each GPU node in the Kubernetes cluster. The MPS daemon provides an independent computing context for each tenant's task on the GPU node, ensuring that the tasks remain independent and isolated while sharing physical GPU resources. During task execution, the MPS daemon dynamically adjusts resource allocation based on the time-slice policy instructions provided by the intelligent scheduling layer, ensuring that each task receives the appropriate computing resources and video memory. By using the MPS daemon, multiple tasks can execute within the same GPU context, reducing frequent context switching between tasks.

[0071] It's important to note that to further optimize video memory utilization, the execution optimization layer introduces a video memory optimizer, which consists of a compression engine and an allocation table. The core function of the video memory optimizer is to monitor video memory usage during task execution and trigger the memory compression mechanism when video memory fragmentation exceeds a certain threshold.

[0072] In one specific embodiment, while executing a target GPU task using a target virtual GPU unit based on target video memory constraints, the computational context corresponding to the target GPU task, and target time slice policy instructions, the fragmentation rate of the video memory is monitored using an allocation table in a preset video memory optimizer. If the fragmentation rate exceeds a fragmentation rate threshold corresponding to the target video memory constraint, a target video memory compression rate is determined, and the target compression engine is used to compress the video memory of the target virtual GPU unit based on the target video memory compression rate, and the allocation table is updated based on the compressed video memory. That is, the compression engine dynamically adjusts the video memory compression ratio based on video memory usage, reducing video memory fragmentation and ensuring that video memory can be efficiently allocated to the currently executing task. The intelligent scheduling layer transmits the target video memory constraint, i.e., the video memory QoS (Quality of Service) constraint, to the video memory optimizer to ensure that the task receives a minimum video memory guarantee, thereby avoiding task performance degradation due to insufficient video memory.

[0073] In another specific embodiment, the target virtual GPU unit is used to execute the target GPU task based on the target memory constraint, the computing context corresponding to the target GPU task, and the target time slice policy instruction. The actual execution time slice of the target GPU task can be monitored, and the target load prediction model can be updated based on the target GPU task and the actual execution time slice. Specifically, through the resource monitoring in the Kubernetes cluster, real-time feedback on the use of GPU resources, including the utilization rate of the computing core and the compression rate of the memory, is fed back to the intelligent scheduling layer, the memory optimizer, and the dynamic time-sharing control layer to help the intelligent scheduling layer and the memory optimizer to perform dynamic resource adjustments. The target load prediction model is updated based on the actual execution time slice of the target GPU task, thereby optimizing the system's resource allocation and task scheduling.

[0074] It should be noted that the workflow for executing the optimization layer can be found in Figure 4 As shown, Figure 4This paper demonstrates how scheduled GPU nodes achieve resource allocation and optimization through the MPS daemon and memory optimizer. Specifically, the MPS daemon provides an independent computing context for each tenant, ensuring task isolation. It also receives time slice policy instructions from the intelligent scheduling layer to dynamically adjust computing resource allocation. The memory optimizer is responsible for dynamic memory allocation and compression management. It records memory usage through an allocation table. When memory fragmentation exceeds a set threshold, it triggers the compression engine to perform memory compression, optimizing memory allocation and utilization and ensuring that idle memory is effectively allocated to executing tasks. The resource monitoring interface provides real-time feedback on compute core utilization and memory compression ratio, providing dynamic resource adjustment data to the intelligent scheduling layer, further optimizing system resource allocation and task execution efficiency. In this way, through the combination of MPS technology and memory optimization mechanisms, the execution optimization layer can achieve efficient GPU resource sharing, task isolation, and memory management in a multi-tenant environment, maximizing GPU resource utilization while ensuring the stability and efficient execution of each task.

[0075] As can be seen from the above, in this embodiment, the target GPU task is first obtained, and the GPU resource requirement information of the target GPU task is determined based on the preset scheduling extender in the Kubernetes cluster, and the resource load of each GPU node in the Kubernetes cluster is evaluated; then the target load prediction model is used to predict the task load curve of the target GPU task, and the target time slice policy instruction corresponding to the target GPU task is generated based on the task load curve; the target load prediction model is a model obtained by training based on the long short-term memory model; then, based on the task load curve, the GPU resource requirement information and the resource load of each GPU node, the target GPU node in the Kubernetes cluster that meets the execution conditions corresponding to the target GPU task is determined; finally, the target virtual GPU unit in the target GPU node is determined, and the target virtual GPU unit is used to execute the target GPU task based on the target time slice policy instruction. As can be seen from the above, in this embodiment, the preset scheduling extender in the Kubernetes cluster is first used to determine the GPU resource requirement information of the target GPU task and evaluate the resource load of each GPU node in the Kubernetes cluster; then, the target load prediction model obtained by training the long short-term memory model is used to predict the task load curve of the target GPU task, and the target time slice policy instruction of the target GPU task is generated according to the task load curve; then, the target GPU node in the Kubernetes cluster that meets the execution conditions of the target GPU task is determined based on the task load curve, the GPU resource requirement information, and the resource load of each GPU node; finally, the target virtual GPU unit in the target GPU node is determined, and the target virtual GPU unit is used to execute the target GPU task based on the target time slice policy instruction. In this way, in this embodiment, by dividing the physical GPU into virtual GPU units in the Kubernetes cluster, and then using the virtual GPU units in the GPU nodes that meet the GPU task execution conditions to execute the GPU task, GPU resources can be finely allocated, avoiding waste and improving resource utilization efficiency. In this way, this embodiment can achieve efficient sharing and fine allocation of GPU resources in a multi-tenant environment, has strong scalability, and meets the needs of GPU virtualization management and resource scheduling in large-scale cloud computing environments.

[0076] In one embodiment, see Figure 5As shown in the figure, the specific workflow of the Kubernetes-based GPU task execution method is as follows: When a new task requests GPU resources, the Scheduler Extender obtains the task's resource requirements and the real-time load information of each GPU node in the Kubernetes cluster, and passes the task load prediction result to the LSTM prediction module. Based on the task load prediction, the Scheduler Extender determines whether the resources of the current GPU node meet the task requirements. If resources are sufficient, the system allocates a virtual GPU unit in the current GPU node to the task and dynamically sets the time slice length. If resources are insufficient, the task is added to the priority queue and waits for scheduling on the next available GPU node. This process is repeated until appropriate resources are allocated for all tasks.

[0077] Accordingly, see Figure 6 As shown, the embodiment of the present application also provides a GPU task execution device based on Kubernetes, which may include:

[0078] A resource evaluation module 11 is configured to obtain a target GPU task, determine the GPU resource requirement information of the target GPU task based on a preset scheduling extender in the Kubernetes cluster, and evaluate the resource load of each GPU node in the Kubernetes cluster;

[0079] A time slice policy instruction generation module 12 is configured to predict a task load curve of the target GPU task using a target load prediction model, and generate a target time slice policy instruction corresponding to the target GPU task based on the task load curve; the target load prediction model is a model trained based on a long short-term memory model;

[0080] A target GPU node determination module 13 is configured to determine a target GPU node in the Kubernetes cluster that meets the execution condition corresponding to the target GPU task based on the task load curve, the GPU resource requirement information, and the resource load of each GPU node;

[0081] The target GPU task execution module 14 is configured to determine a target virtual GPU unit in the target GPU node, and utilize the target virtual GPU unit to execute the target GPU task based on the target time slice policy instruction.

[0082] As can be seen from the above, in this application, the target GPU task is first obtained, and the GPU resource requirement information of the target GPU task is determined based on the preset scheduling extender in the Kubernetes cluster, and the resource load of each GPU node in the Kubernetes cluster is evaluated; then the target load prediction model is used to predict the task load curve of the target GPU task, and the target time slice policy instruction corresponding to the target GPU task is generated based on the task load curve; the target load prediction model is a model obtained by training based on the long short-term memory model; then, based on the task load curve, the GPU resource requirement information and the resource load of each GPU node, the target GPU node in the Kubernetes cluster that meets the execution conditions corresponding to the target GPU task is determined; finally, the target virtual GPU unit in the target GPU node is determined, and the target virtual GPU unit is used to execute the target GPU task based on the target time slice policy instruction. As can be seen from the above, in this application, the preset scheduling extender in the Kubernetes cluster is first used to determine the GPU resource requirement information of the target GPU task and evaluate the resource load of each GPU node in the Kubernetes cluster; then, the target load prediction model obtained by training the long short-term memory model is used to predict the task load curve of the target GPU task, and the target time slice policy instruction of the target GPU task is generated according to the task load curve; then, the target GPU node in the Kubernetes cluster that meets the execution conditions of the target GPU task is determined according to the task load curve, GPU resource requirement information, and the resource load of each GPU node; finally, the target virtual GPU unit in the target GPU node is determined, and the target virtual GPU unit is used to execute the target GPU task based on the target time slice policy instruction. In this way, the present application can refine the allocation of GPU resources in the Kubernetes cluster by dividing the physical GPU into virtual GPU units, and then using the virtual GPU units in the GPU nodes that meet the GPU task execution conditions to execute GPU tasks, thereby avoiding waste and improving resource utilization efficiency. In this way, the present application can achieve efficient sharing and refined allocation of GPU resources in a multi-tenant environment, has strong scalability, and meets the needs of GPU virtualization management and resource scheduling in large-scale cloud computing environments.

[0083] In some specific implementations, the Kubernetes-based GPU task execution device may further include:

[0084] A virtual GPU unit partitioning module is configured to deploy a preset device plug-in on each GPU node in the Kubernetes cluster, use the preset device plug-in to partition each GPU node into a plurality of virtual GPU units based on a preset business scenario, and register each virtual GPU unit with the Kubernetes cluster using a target remote procedure call framework;

[0085] The target GPU resource management table determination module is configured to define the virtual GPU resource attributes and operating status of each virtual GPU unit in each GPU node in the Kubernetes cluster through a preset resource management mechanism to obtain a corresponding target GPU resource management table; the virtual GPU resource attributes include computing cores and video memory, and the operating status includes idle and occupied.

[0086] In some specific implementations, the resource evaluation module 11 may include:

[0087] a target idle computing core number determining unit, configured to determine the number of first virtual GPU units in an idle working state in each GPU node based on the preset scheduling extender and the target GPU resource management table in the Kubernetes cluster, and determine the target idle computing core number and first idle video memory corresponding to the first virtual GPU unit;

[0088] a target idle video memory determining unit, configured to determine the second idle video memory corresponding to each GPU node based on an allocation table in a preset video memory optimizer, and determine a corresponding target idle video memory based on the first idle video memory and the second idle video memory;

[0089] A resource load evaluation unit is used to evaluate the resource load of each GPU node in the Kubernetes cluster based on the target number of idle computing cores, the target number of idle video memory, and the number of the first virtual GPU units.

[0090] In some specific implementations, the Kubernetes-based GPU task execution device may further include:

[0091] A historical task load data acquisition module is used to determine the task type corresponding to the target GPU task, obtain several historical GPU tasks of the same task type, and obtain historical task load data corresponding to each of the historical GPU tasks; the task types include inference tasks and training tasks;

[0092] a target load prediction model determination module, configured to use the long short-term memory model to learn and analyze the historical task load data corresponding to each of the historical GPU tasks to obtain the target load prediction model;

[0093] Accordingly, the time slice policy instruction generating module 12 may include:

[0094] a task load curve determining unit, configured to input the GPU resource requirement information corresponding to the target GPU task into the target load prediction model, and output the task load curve corresponding to the target GPU task using the target load prediction model;

[0095] A time slice policy instruction generating unit is configured to generate the target time slice policy instruction corresponding to the target GPU task based on the task type corresponding to the target GPU task and the task load curve.

[0096] In some specific implementations, the Kubernetes-based GPU task execution device may further include:

[0097] The GPU task execution module to be executed is used to determine the target priority corresponding to the target GPU task based on the task type corresponding to the target GPU task if it is determined based on the task load curve, the GPU resource demand information and the resource load of each GPU node that there is no target GPU node in the Kubernetes cluster that meets the execution conditions corresponding to the target GPU task, and add the target GPU task to the target GPU task queue to be executed based on the target priority, so as to execute each GPU task to be executed based on the priority corresponding to each GPU task to be executed in the target GPU task queue to be executed.

[0098] In some specific implementations, the Kubernetes-based GPU task execution device may further include:

[0099] A multi-process service daemon deployment module is used to deploy a multi-process service daemon on each GPU node in the Kubernetes cluster;

[0100] A computing context creation module is used to create a computing context corresponding to the target GPU task by using the multi-process service daemon on the target GPU node, and obtain a target time slice policy instruction corresponding to the target GPU task by using the multi-process service daemon;

[0101] Accordingly, the target GPU task execution module 14 may include:

[0102] A target video memory constraint determination submodule, configured to bind the target virtual GPU unit to the target GPU task and determine a target video memory constraint corresponding to the target GPU task;

[0103] The target GPU task execution submodule is used to utilize the target virtual GPU unit to execute the target GPU task based on the target video memory constraint, the computing context corresponding to the target GPU task and the target time slice policy instruction.

[0104] In some specific implementations, the target GPU task execution submodule may further include:

[0105] a video memory compression unit, configured to monitor a fragmentation rate of the video memory using an allocation table in a preset video memory optimizer, determine a target video memory compression rate if the fragmentation rate exceeds a fragmentation rate threshold corresponding to the target video memory constraint, compress the video memory of the target virtual GPU unit based on the target video memory compression rate using a target compression engine, and update the allocation table based on the compressed video memory;

[0106] The target load prediction model updating unit is configured to monitor the actual execution time slice of the target GPU task and update the target load prediction model based on the target GPU task and the actual execution time slice.

[0107] Furthermore, the embodiment of the present application also discloses an electronic device, Figure 7 This is a structural diagram of an electronic device 20 shown according to an exemplary embodiment. The content in the diagram cannot be considered as any limitation on the scope of use of this application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 is used to store a computer program, which is loaded and executed by the processor 21 to implement the relevant steps in the Kubernetes-based GPU task execution method disclosed in any of the aforementioned embodiments. In addition, the electronic device 20 in this embodiment may specifically be an electronic computer.

[0108] In this embodiment, the power supply 23 is used to provide operating voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and the external device. The communication protocol it follows is any communication protocol that can be applied to the technical solution of this application and is not specifically limited here; the input and output interface 25 is used to obtain external input data or output data to the outside world. Its specific interface type can be selected according to specific application needs and is not specifically limited here.

[0109] In addition, the memory 22, as a carrier for resource storage, can be a read-only memory, random access memory, disk or CD, etc. The resources stored thereon can include an operating system 221, a computer program 222, etc., and the storage method can be temporary storage or permanent storage.

[0110] The operating system 221 is used to manage and control the hardware devices on the electronic device 20 and the computer program 222, and can be Windows Server, Netware, Unix, Linux, etc. In addition to including a computer program capable of implementing the Kubernetes-based GPU task execution method performed by the electronic device 20 disclosed in any of the aforementioned embodiments, the computer program 222 may further include a computer program capable of performing other specific tasks.

[0111] Furthermore, this application also discloses a computer-readable storage medium for storing a computer program; wherein, when executed by a processor, the computer program implements the aforementioned Kubernetes-based GPU task execution method. The specific steps of this method can be referred to the corresponding content disclosed in the aforementioned embodiments and will not be repeated here.

[0112] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from the other embodiments. Reference can be made to the descriptions of the identical or similar parts between the various embodiments. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple, and the relevant parts can be referred to the descriptions of the methods.

[0113] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0114] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.

[0115] Finally, it should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.

[0116] The above is a detailed introduction to the technical solution provided by the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea. At the same time, for those skilled in the art, according to the ideas of the present application, there may be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.

Claims

1. A GPU task execution method based on Kubernetes, characterized in that: include: Obtain a target GPU task, determine the GPU resource requirement information of the target GPU task based on a preset scheduling extender in the Kubernetes cluster, and evaluate the resource load of each GPU node in the Kubernetes cluster; A target load prediction model is used to predict a task load curve of the target GPU task, and a target time slice policy instruction corresponding to the target GPU task is generated based on the task load curve; the target load prediction model is a model obtained by training based on a long short-term memory model; Determine a target GPU node in the Kubernetes cluster that meets the execution condition corresponding to the target GPU task based on the task load curve, the GPU resource requirement information, and the resource load of each GPU node; A target virtual GPU unit in the target GPU node is determined, and the target virtual GPU unit is used to execute the target GPU task based on the target time slice policy instruction.

2. The Kubernetes-based GPU task execution method according to claim 1, characterized in that: Also includes: Deploy a preset device plug-in on each GPU node in the Kubernetes cluster, divide each GPU node into a plurality of virtual GPU units based on a preset business scenario using the preset device plug-in, and register each virtual GPU unit with the Kubernetes cluster using a target remote procedure call framework; The virtual GPU resource attributes and working status of each virtual GPU unit in each GPU node are defined in the Kubernetes cluster through a preset resource management mechanism to obtain a corresponding target GPU resource management table; the virtual GPU resource attributes include computing cores and video memory, and the working status includes idle and occupied.

3. The Kubernetes-based GPU task execution method according to claim 2, characterized in that: Evaluate the resource load of each GPU node in the Kubernetes cluster based on a preset scheduling extender in the Kubernetes cluster, including: Determine the number of first virtual GPU units in idle working state in each GPU node based on the preset scheduling extender and the target GPU resource management table in the Kubernetes cluster, and determine the target number of idle computing cores and first idle video memory corresponding to the first virtual GPU unit; Determine the second idle video memory corresponding to each GPU node based on an allocation table in a preset video memory optimizer, and determine a corresponding target idle video memory based on the first idle video memory and the second idle video memory; The resource load of each GPU node in the Kubernetes cluster is evaluated based on the target number of idle computing cores, the target idle video memory, and the number of the first virtual GPU units.

4. The GPU task execution method based on Kubernetes according to claim 1, characterized in that Before predicting the task load curve of the target GPU task using the target load prediction model, the method further includes: Determine a task type corresponding to the target GPU task, obtain several historical GPU tasks of the same task type, and obtain historical task load data corresponding to each of the historical GPU tasks; the task types include inference tasks and training tasks; Using the long short-term memory model to learn and analyze the historical task load data corresponding to each of the historical GPU tasks to obtain the target load prediction model; Accordingly, the method of predicting the task load curve of the target GPU task using the target load prediction model and generating the target time slice policy instruction corresponding to the target GPU task based on the task load curve includes: Inputting the GPU resource requirement information corresponding to the target GPU task into the target load prediction model, and outputting the task load curve corresponding to the target GPU task using the target load prediction model; The target time slice policy instruction corresponding to the target GPU task is generated based on the task type corresponding to the target GPU task and the task load curve.

5. The Kubernetes-based GPU task execution method according to claim 1, wherein: Also includes: If it is determined based on the task load curve, the GPU resource requirement information and the resource load of each GPU node that there is no target GPU node in the Kubernetes cluster that meets the execution conditions corresponding to the target GPU task, the target priority corresponding to the target GPU task is determined based on the task type corresponding to the target GPU task, and the target GPU task is added to the target to-be-executed GPU task queue based on the target priority, so that each to-be-executed GPU task is executed based on the priority corresponding to each to-be-executed GPU task in the target to-be-executed GPU task queue.

6. The Kubernetes-based GPU task execution method according to any one of claims 1 to 5, characterized in that: Before executing the target GPU task based on the target time slice policy instruction using the target virtual GPU unit, the method further includes: Deploy a multi-process service daemon on each GPU node in the Kubernetes cluster; Creating a computing context corresponding to the target GPU task by using the multi-process service daemon on the target GPU node, and obtaining a target time slice policy instruction corresponding to the target GPU task by using the multi-process service daemon; Accordingly, the using the target virtual GPU unit to execute the target GPU task based on the target time slice policy instruction includes: Bind the target virtual GPU unit to the target GPU task, and determine the target video memory constraint condition corresponding to the target GPU task; The target GPU task is executed using the target virtual GPU unit based on the target video memory constraint, the computing context corresponding to the target GPU task, and the target time slice policy instruction.

7. The Kubernetes-based GPU task execution method according to claim 6, characterized in that: The process of executing the target GPU task using the target virtual GPU unit based on the target video memory constraint, the computing context corresponding to the target GPU task, and the target time slice policy instruction further includes: monitoring a fragmentation rate of the video memory using an allocation table in a preset video memory optimizer; determining a target video memory compression rate if the fragmentation rate exceeds a fragmentation rate threshold corresponding to the target video memory constraint; compressing the video memory of the target virtual GPU unit based on the target video memory compression rate using a target compression engine; and updating the allocation table based on the compressed video memory; An actual execution time slice of the target GPU task is monitored, and the target load prediction model is updated based on the target GPU task and the actual execution time slice.

8. A GPU task execution device based on Kubernetes, characterized in that: include: A resource assessment module is used to obtain a target GPU task, determine the GPU resource requirement information of the target GPU task based on a preset scheduling extender in the Kubernetes cluster, and assess the resource load of each GPU node in the Kubernetes cluster; A time slice policy instruction generation module is configured to predict a task load curve of the target GPU task using a target load prediction model, and generate a target time slice policy instruction corresponding to the target GPU task based on the task load curve; the target load prediction model is a model trained based on a long short-term memory model; A target GPU node determination module is configured to determine a target GPU node in the Kubernetes cluster that meets the execution conditions corresponding to the target GPU task based on the task load curve, the GPU resource requirement information, and the resource load of each GPU node; The target GPU task execution module is configured to determine a target virtual GPU unit in the target GPU node, and utilize the target virtual GPU unit to execute the target GPU task based on the target time slice policy instruction.

9. An electronic device, characterized in that: The electronic device includes a processor and a memory; wherein the memory is used to store a computer program, and the computer program is loaded and executed by the processor to implement the Kubernetes-based GPU task execution method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that Used to store a computer program, which, when executed by a processor, implements the Kubernetes-based GPU task execution method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Video memory allocation method and device and storage medium

    CN121144052A

  • Cluster computing resource isolation method and device, electronic equipment and storage medium

    CN121833262A