Computing power resource virtualization isolation and multi-level scheduling method and system in containerized environment

By deploying indicator collectors and custom schedulers in the Kubernetes cluster and combining them with vGPU technology, we achieve multi-level scheduling and virtualization isolation of computing resources in a containerized environment, solving the problems of unreasonable resource allocation and task failure, and improving resource utilization and task execution efficiency.

CN120704888APending Publication Date: 2025-09-26HANGZHOU INTERNATIONAL INNOVATION INSTITUTE OF BEIHANG UNIVERSITY

Patent Information

Application Number
CN202510843724.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-23
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

Existing technologies make it difficult to effectively manage and schedule heterogeneous computing resources in a containerized environment, resulting in irrational resource allocation, idle waste, and task failure. In addition, the lack of virtualization isolation of computing resources makes it impossible to meet the requirements for stable task operation in a multi-tenant environment.

Method used

A method of virtualization isolation and multi-level scheduling of computing resources in a containerized environment is adopted. By deploying indicator collectors and custom schedulers on the Kubernetes cluster, resource utilization is monitored in real time. vGPU technology is used to achieve fine-grained allocation and scheduling of computing resources. Combined with a multi-level scheduling algorithm, the most suitable working nodes and task bindings are selected to achieve efficient resource utilization and isolation.

Benefits of technology

It improves resource utilization and load balancing, enhances task execution efficiency, realizes secure sharing of physical and virtual resources, reduces system deployment difficulty and cost, and enhances compatibility with Kubernetes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120704888A_ABST
    Figure CN120704888A_ABST
Patent Text Reader

Abstract

The invention relates to a computing power resource virtualization isolation and multi-level scheduling method and system in a containerization environment, and the method comprises the steps: deploying an index collector on each working node of a Kubernetes cluster, and collecting the node computing resource utilization condition in real time; when a calculation task request submitted by a user is received, the Kubernetes API server performs field legality verification on the configuration file submitted by the user and stores the field legality verification; executing a computing power resource multi-level scheduling algorithm according to a configuration file submitted by a user and a resource utilization condition, and finding a working node with the highest adaptation degree as an optimal node; the user-defined scheduler sends an instruction to the optimal node, requires the optimal node to create a vGPU and a Pod, and mounts the vGPU to the corresponding Pod to run a calculation task; and recording vGPU distribution information of the optimal node, storing and updating the cluster. According to the invention, computing power resource sharing and isolation are realized, and the resource utilization rate is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computing power resource scheduling, and in particular to a method and system for virtualization isolation and multi-level scheduling of computing power resources in a containerized environment. Background Art

[0002] Containerization, due to its lightweight, efficient, and portable nature, has gradually become a mainstream technology in the cloud computing field. With the rapid development of artificial intelligence and high-performance computing, GPU computing resources are widely used in cloud computing environments. Efficiently managing and scheduling computing resources, especially GPU computing resources, in containerized environments has become a key issue. Kubernetes, as a mainstream container orchestration system, has become the foundational platform for resource scheduling in cloud-native environments. However, existing Kubernetes-based computing resource scheduling solutions still face many technical bottlenecks.

[0003] Traditional computing resource management methods struggle to accurately match different types of tasks, resulting in irrational resource allocation, wasted idle resources, and the inability to execute other tasks in a timely manner due to insufficient resources. Different tasks have varying requirements for computing resources, necessitating an urgent solution for effectively scheduling heterogeneous computing resources like CPUs and GPUs to meet task requirements and improve resource utilization. The default Kubernetes scheduling policy only supports the management and allocation of CPU and memory resources, not the identification and management of GPU resources. Scheduling decisions are based solely on static resource status, ignoring the real-time load status of nodes. This can easily lead to imbalanced resource allocation, with overloaded "hot nodes" and idle "cold nodes." Furthermore, the policy lacks a quantitative assessment of task execution efficiency, making it impossible to optimize scheduling based on the characteristics of tasks such as compute-intensive or IO-intensive tasks.

[0004] In a multi-tenant containerized environment, when multiple tasks share the same physical resources, resource contention is prone to occur, leading to performance degradation or even task failure. Achieving virtualized isolation of computing resources and ensuring stable task operation is a critical issue in resource management in containerized environments. Hardware vendors' default GPU device plug-in mechanism only supports allocation of entire GPU cards and fails to implement fine-grained partitioning of physical GPUs. This results in tasks requiring the entire GPU card even when only a portion of the computing power is required, resulting in significant waste of computing power. Existing third-party GPU virtualization solutions often rely on kernel-level driver modifications or specialized hardware, lacking compatibility with the Kubernetes ecosystem and failing to adapt to the unified management requirements of heterogeneous GPU devices in multi-cloud environments. Despite industry efforts to optimize through GPU device plug-in extensions and scheduler customization, efficient computing resource management remains impeded due to fundamental flaws such as the decoupling of resource scheduling from virtualization and a lack of a global optimization perspective. Therefore, a solution that deeply integrates vGPU virtualization technology with resource scheduling algorithms is urgently needed to build a comprehensive computing resource isolation and scheduling technology system to improve resource utilization and task execution efficiency. Summary of the Invention

[0005] The technology of the present invention solves the problem: Overcoming the shortcomings of the existing technology, it provides a method and system for virtualization isolation and multi-level scheduling of computing power resources in a containerized environment, which fully schedules different types of computing power and effectively improves resource utilization and load balancing.

[0006] The present invention adopts the following technical solutions:

[0007] A method for virtualization isolation and multi-level scheduling of computing resources in a containerized environment, comprising the following steps:

[0008] Step S110: deploy an indicator collector on each working node of the Kubernetes cluster to collect the node computing resource utilization in real time;

[0009] Step S120: Upon receiving a computing task request submitted by a user, the Kubernetes API server verifies the legitimacy of the fields in the configuration file submitted by the user and stores the result.

[0010] Step S130: Deploy a custom scheduler on the Kubernetes control plane, execute a multi-level computing resource scheduling algorithm based on the user-submitted configuration file and resource utilization, and find the working node with the highest adaptability as the optimal node;

[0011] Step S140: The custom scheduler notifies the optimal node to create a vGPU and a Pod, and mounts the vGPU to the corresponding Pod to run the computing task;

[0012] Step S150: Record the vGPU allocation information of the optimal node, store it, and update the cluster.

[0013] A computing resource virtualization isolation and multi-level scheduling system in a containerized environment includes: a resource monitoring module, a task request processing module, a custom scheduler module, and a vGPU module; wherein:

[0014] The resource monitoring module is responsible for real-time monitoring of the resource utilization of the working nodes of the Kubernetes cluster and providing the monitoring data to the custom scheduler module for screening the adaptability of nodes and computing nodes to tasks;

[0015] The task request processing module is responsible for receiving task requests submitted by users by creating Pod configuration files, verifying the rationality of the configuration file fields and storing them, and providing the task resource requirements to the custom scheduler module for screening the nodes and computing nodes for their suitability with the task;

[0016] The custom scheduler module first obtains the node's real-time resource utilization indicators from the resource monitoring module. Based on the task resource requirements provided by the task request processing module, it filters out nodes whose current resource availability is less than the task resource requirements. It then uses a custom scheduling algorithm to calculate the execution efficiency and load balancing fitness between each node and task, performing a weighted summation to obtain the overall fitness between the node and task. It then selects the node with the highest overall fitness and binds it to the task. The task-node binding information is then sent to the vGPU management module.

[0017] After the vGPU management module receives the task-node binding information sent by the custom scheduler, the vGPU device plug-in manager in the Kubelet on the working node applies to the vGPU device plug-in to create a vGPU. The vGPU device plug-in creates a vGPU in the managed GPU resources and returns the ID. The Kubelet creates a Pod corresponding to the task, mounts the Pod on the created vGPU, starts the container to run the computing task, and completes the scheduling of computing resources. At the same time, the vGPU device plug-in reports the created vGPU allocation information to the Kubelet. The Kubelet saves the vGPU allocation information in the Pod cache and further reports it to the control plane for storage. When the corresponding Pod ends or the status is abnormal, the corresponding vGPU is destroyed to achieve resource recovery.

[0018] The beneficial effects of the present invention compared with the prior art are:

[0019] (1) This paper quantifies the adaptability of work nodes and tasks from the perspectives of execution efficiency and load balancing by comprehensively considering task resource requirements and resource utilization. It then proposes a multi-level scheduling algorithm to select the most suitable work nodes and task bindings, thereby improving resource utilization.

[0020] (2) The present invention establishes a mapping relationship between physical resources and virtual resources through virtual computing power devices and service programs, realizes the isolation of computing power resources at the logical level, enables different containers to safely share the same physical GPU, and improves the utilization rate of the GPU.

[0021] (3) This invention designs a corresponding system based on the containerized environment and Kubernetes architecture, targeting the deployment characteristics of existing cloud computing environments. The system is compatible with native Kubernetes, can seamlessly connect and fully utilize its core functions; the system deployment process is simple and fast, with low deployment difficulty and cost; the system operation process is clear and interactive, which can effectively improve the user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 This is a flow chart of a method for virtualization isolation and multi-level scheduling of computing resources in a containerized environment according to the present invention;

[0023] Figure 2 The resource monitoring module of the present invention;

[0024] Figure 3 It is a task request processing module of the present invention;

[0025] Figure 4 The custom scheduler module of the present invention;

[0026] Figure 5 This is the vGPU management module of the present invention. DETAILED DESCRIPTION

[0027] Exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments described herein. Rather, these embodiments are provided to enable a more thorough understanding of the present disclosure and to fully convey the scope of the present disclosure to those skilled in the art.

[0028] Figure 1 This is a flow chart of a method for virtualization isolation and multi-level scheduling of computing resources in a containerized environment. Figure 1 As shown, the method specifically includes the following steps:

[0029] Step S110: deploy an indicator collector on each working node of the Kubernetes cluster to collect the node computing resource utilization in real time;

[0030] Step S120: Upon receiving a computing task request submitted by a user, the Kubernetes API server verifies the legitimacy of the fields in the configuration file submitted by the user and stores the result.

[0031] Step S130: deploy a custom scheduler on the Kubernetes control plane, execute a multi-level scheduling algorithm for computing resources, and find the working node with the highest fitness as the optimal node;

[0032] Step S140: The custom scheduler sends instructions to the optimal node, requesting it to create a vGPU and a Pod, and mount the vGPU to the corresponding Pod to run the computing task;

[0033] Step S150: Record the vGPU allocation information of the optimal node, store it, and update the cluster.

[0034] In step S110, each working node of the Kubernetes cluster includes:

[0035] (1) Physical GPU resources refer to the physical graphics processor hardware devices installed on the worker nodes, which provide raw computing power support;

[0036] (2) vGPU device plug-in, which is responsible for registering and managing physical GPU resources with Kubelet, and performing vGPU segmentation, status monitoring, and device allocation operations. It can split a physical GPU into multiple vGPUs and limit the video memory and computing units so that different containers can safely share the same physical GPU;

[0037] (3) Virtual GPU resources (vGPU) refer to virtualized computing units formed by dividing physical GPUs through vGPU device plug-ins. Each vGPU has independent video memory isolation domain and computing core quota;

[0038] (4) Pod, which refers to the smallest task unit scheduled by Kubernetes and contains one or more containers that share a computing environment;

[0039] (5) Kubelet, which runs on each working node of the cluster and is responsible for managing the life cycle of the Pod on the running node, resource monitoring, node status checking, container start and stop, etc.

[0040] In step S110, the indicator collector collects the node computing resource utilization in real time, specifically:

[0041] (1) The indicator collectors include: general computing power indicator collector, intelligent computing power indicator collector and Prometheus service;

[0042] (2) The general computing power indicator collector is deployed on the node where the CPU exists to collect the utilization of CPU and memory related indicators;

[0043] (3) The intelligent computing power indicator collector is deployed on the node with GPU, monitoring the static performance indicators such as the number of GPUs, video memory capacity, GPU computing resource utilization, and video memory utilization of the node;

[0044] (4) The Prometheus service pulls data from the general computing power indicator collector and the intelligent computing power indicator collector every 5 seconds and stores it in the local time series database.

[0045] In step S120, when a computing task request submitted by a user is received, the configuration file submitted by the user is verified for field legitimacy and stored, specifically:

[0046] (1) The user submits a Pod configuration file in YAML format through an API request, which contains a resource requirement statement. The resource requirement statement is the original form in the configuration file, which is parsed and abstracted into task resource requirements, including resource requests (minimum amount) and resource limits (maximum amount).

[0047] (2) The Kubernetes API server verifies the legality of the name, type, format, and value range of the fields in the resource requirement declaration based on the custom resource definition (CRD);

[0048] (3) After verification, the Kubernetes API server stores the Pod configuration in etcd.

[0049] In step S130, the Kubernetes control plane includes:

[0050] (1) The Kubernetes API server, as the only entry point to the Kubernetes API, is the hub for interaction between various components and is responsible for handling all communication and data interaction requests from users or components within the cluster;

[0051] (2) Custom scheduler: When a new Pod issues a scheduling request, the custom scheduler will schedule the Pod to the most suitable node based on the resource amount requested by the Pod and the cluster resource utilization according to the custom scheduling algorithm;

[0052] (3) etcd refers to the Kubernetes backend persistent storage database that stores the status data of the entire cluster.

[0053] In step S130, the multi-level scheduling algorithm for computing resources is implemented as follows:

[0054] (1) The custom scheduler obtains the Pod configuration file by listening to the Kubernetes API server, parses the resource request (the minimum amount of resources that the Pod needs to ensure) and resource limit (the maximum amount of resources allowed to the Pod) to obtain the task resource requirements, including the requirements for CPU, memory, number of GPUs, and GPU memory resources;

[0055] (2) The custom scheduler queries the Prometheus service for resource utilization indicators and receives the monitoring data returned by the Prometheus service;

[0056] (3) At the first level, the custom scheduler filters out unqualified worker nodes from the computing resource cluster based on Pod resource requests and node resource utilization. Specifically, when a task applies for GPU computing resources, the custom scheduler filters out all worker nodes that do not contain GPUs. When the Pod resource request is greater than the remaining computing resources of the node, the custom scheduler filters out this node.

[0057] (4) At the second level, for nodes that meet the conditions, the multi-level scheduling algorithm of computing power resources calculates the fitness of different nodes and tasks, finds the optimal node and binds the task; the fitness of the node and the task is composed of the weighted sum of the execution efficiency fitness and the load balancing fitness. The execution efficiency fitness reflects the execution time competitiveness of the task at the node, and the load balancing fitness represents the contribution of task allocation to the global resource load balancing. The weight coefficient is dynamically adjusted according to the real-time status of the task and the node. When the task type is sensitive to the execution time, the execution efficiency weight is increased, and when there is a significant uneven distribution of resources in the system, the load balancing weight is increased.

[0058] In this multi-level computing resource scheduling algorithm, the execution efficiency adaptation and load balancing adaptation are implemented as follows:

[0059] (1) Execution efficiency adaptability: Based on the task's demand for CPU or GPU resources, the time required for the node to execute the task is calculated, and a semi-normalized evaluation model is constructed by comparing the target node's execution time with the longest execution time of the entire cluster. The larger the value, the better the execution efficiency of the node compared to the cluster as a whole.

[0060] (2) Load balancing adaptability uses the standard deviation of resource utilization of all nodes in the cluster to quantify the degree of global load balancing. The node resource utilization is composed of the weighted sum of CPU and GPU resource utilization. The weight coefficient can be dynamically adjusted according to the specific scenario. The larger the value, the more effective the task allocation is in improving the overall load balancing status of the system.

[0061] In step S140, the optimal node creates a vGPU and a Pod and completes the mounting, specifically:

[0062] (1) The custom scheduler sends the task-node binding information to the optimal node;

[0063] (2) The vGPU device plug-in manager in the Kubelet of the optimal node requests the vGPU device plug-in to allocate a vGPU;

[0064] (3) The vGPU device plug-in creates a vGPU and reports the corresponding vGPU allocation information to Kubelet, including the vGPU ID;

[0065] (4) Kubelet creates a Pod and mounts it to the corresponding vGPU to complete the scheduling. It also saves the vGPU allocation information to the node cache and reports it to the Kubernetes control plane.

[0066] (5) The Kubernetes API server stores the received vGPU allocation information in etcd.

[0067] The above method may further comprise the steps of:

[0068] S160, the Kubernetes control plane monitors and obtains Pod destruction events and health status;

[0069] S170: The vGPU device plug-in destroys the vGPU corresponding to the Pod that is out of lifecycle or unhealthy to reclaim GPU resources.

[0070] like Figure 1 As shown, a specific example of the above method steps can be as follows:

[0071] (1) The user defines the Pod configuration file and submits a task request to the Kubernetes API server;

[0072] (2) The Kubernetes API server receives the Pod configuration file and checks its syntax. If the syntax is correct, the Kubernetes API server stores the Pod configuration file in the etcd database.

[0073] (3) The custom scheduler obtains the Pod configuration file by listening to the Kubernetes API server, parses the resource request (the minimum amount of resources that the Pod needs to ensure) and resource limit (the maximum amount of resources allowed to the Pod) in the Pod configuration file, and obtains the task resource requirements, including the requirements for CPU, memory, number of GPUs, and GPU memory resources;

[0074] (4) The custom scheduler queries the Prometheus service for resource utilization indicators and receives the monitoring data returned by the Prometheus service;

[0075] (5) The custom scheduler executes a multi-level scheduling algorithm and returns the node ID that best matches the task to the Kubernetes API server;

[0076] (6) The Kubernetes API server receives the task-node binding information and stores it in etcd;

[0077] (7) The Kubernetes API server transmits the task-node binding information to the Kubelet of the corresponding node;

[0078] (8) The vGPU device plug-in manager in Kubelet requests the vGPU device plug-in to allocate the vGPU required by the corresponding Pod;

[0079] (9) The vGPU device plug-in creates a vGPU;

[0080] (10) The vGPU device plug-in reports the vGPU allocation information, including the vGPU ID, to the vGPU device plug-in manager;

[0081] (11) Kubelet creates a Pod and mounts the corresponding vGPU resources to it, and starts the container;

[0082] (12) Kubelet retains the vGPU allocation information in the node cache in the form of annotations and further reports it to the Kubernetes API server;

[0083] (13) The Kubernetes API server receives the vGPU allocation information and stores it in etcd.

[0084] like Figures 2 to 5 As shown in FIG, the computing power resource virtualization isolation and multi-level scheduling system in the containerized environment of the present invention includes a resource monitoring module, a task request processing module, a custom scheduler module, and a vGPU management module, wherein:

[0085] The resource monitoring module is responsible for real-time monitoring of the resource utilization of the working nodes of the Kubernetes cluster and providing the monitoring data to the custom scheduler module for screening the adaptability of nodes and computing nodes to tasks;

[0086] The task request processing module is responsible for receiving task requests submitted by users by creating Pod configuration files, verifying the rationality of the configuration file fields and storing them, and providing the task resource requirements to the custom scheduler module for screening the nodes and computing nodes for their suitability with the task;

[0087] The custom scheduler module first obtains the node's real-time resource utilization indicators from the resource monitoring module. Based on the task resource requirements provided by the task request processing module, it filters out nodes whose current resource availability is less than the task resource requirements. It then uses a custom scheduling algorithm to calculate the execution efficiency and load balancing fitness between each node and task, and performs a weighted summation to obtain the overall fitness between the node and task. It then selects the node with the highest overall fitness and binds it to the task. The task-node binding information is then sent to the vGPU management module.

[0088] After the vGPU management module receives the task-node binding information from the custom scheduler, the vGPU device plug-in manager in the Kubelet on the worker node requests the creation of a vGPU from the vGPU device plug-in. The vGPU device plug-in creates the vGPU in the managed GPU resources and returns the ID. The Kubelet then creates the Pod corresponding to the task, mounts the Pod on the created vGPU, starts the container to run the computing task, and completes the scheduling of computing resources. At the same time, the vGPU device plug-in reports the created vGPU allocation information to the Kubelet, which saves the vGPU allocation information in the Pod cache and further reports it to the control plane for storage. When the corresponding Pod ends or the status is abnormal, the corresponding vGPU is destroyed to achieve resource recovery.

[0089] The specific implementation process of each module is as follows:

[0090] 1. If Figure 2 As shown, the implementation of the resource monitoring module includes:

[0091] (1) The general computing power indicator collector is deployed on the node with CPU to monitor the utilization of CPU and memory related indicators;

[0092] (2) The intelligent computing power indicator collector is deployed on the node with GPU, monitoring the static performance indicators such as the number of GPUs, video memory capacity, GPU computing resource utilization, and video memory utilization of the node;

[0093] (3) The Prometheus service pulls data from the general computing power indicator collector and the intelligent computing power indicator collector every 5 seconds and stores it in the local time series database;

[0094] (4) The custom scheduler queries the Prometheus service for resource utilization through the API.

[0095] 2. If Figure 3 As shown, the implementation of the task request processing module includes:

[0096] (1) The user submits a Pod configuration file in YAML format through an API request, which contains a resource requirement statement;

[0097] (2) The Kubernetes API server verifies the legality of the name, type, format, and value range of the fields in the resource requirement declaration based on the custom resource definition (CRD);

[0098] (3) After verification, the Kubernetes API server stores the Pod configuration in etcd.

[0099] 3. If Figure 4 As shown, the implementation of the custom scheduler module includes:

[0100] (1) The custom scheduler obtains the Pod configuration file by listening to the Kubernetes API server, parses the resource request (the minimum amount of resources that the Pod needs to ensure) and resource limit (the maximum amount of resources allowed to the Pod) in the Pod configuration file, and obtains the task resource requirements, including the requirements for CPU, memory, number of GPUs, and GPU memory resources;

[0101] (2) The custom scheduler queries the Prometheus service for resource utilization indicators and receives the monitoring data returned by the Prometheus service;

[0102] (3) Based on the Pod resource request and resource utilization, filter out the worker nodes whose current idle resources are less than the number of resources required by the task;

[0103] (4) Traverse the remaining working nodes, and for each node, assign the computing task to the node Execution efficiency adaptability: ,in, , to assign tasks to nodes The execution time, , represents a set of task types, Indicates the Pod resource request for the corresponding type of computing power. Represents a working node The computing power resources of the corresponding chip type that can be provided. The resource request and computing power resource units are FLOPs / GFLOPs. , represents the longest execution time of the task in all working nodes, The value range is [0.5, 1). The larger the value, the more working nodes The smaller the relative time cost of processing the task;

[0104] (5) Computation tasks are assigned to nodes Load balancing adaptability: ,in, , represents global load balancing, estimated using standard deviation, Indicates the total number of nodes, , represents the average resource usage of all nodes, node The overall resource utilization rate is , Indicates the Pod resource request for CPU / GPU type computing power. Indicates the The total amount of CPU / GPU computing power on the worker nodes, Indicates the current The CPU / GPU computing power utilization on the working nodes and the comprehensive resource utilization of other nodes are: , 、 Indicates the importance coefficient of CPU and GPU load balancing, satisfying , can be dynamically adjusted according to the real-time situation of tasks and nodes. When the CPU node load balancing degree is low, set a larger When the GPU node load balancing is low, set a larger , The range is (0.5, 1], and a larger value indicates a higher degree of global load balancing.

[0105] (6) Calculate the adaptability of nodes and tasks through weighted summation: ,in Represents a working node Suitability to the task, Indicates the suitability of execution efficiency. Indicates the load balancing adaptability. 、 Represents the importance weight coefficient, satisfying , can be dynamically adjusted according to the real-time status of tasks and nodes. When the computing time efficiency is relatively low, set a larger , when the node load balancing degree is low, set a larger ;

[0106] (7) After the traversal is completed, the working node with the largest node-task adaptability is obtained, and it is determined as the optimal node and bound to the task;

[0107] (8) The task-node binding information is stored in etcd through the Kubernetes API server.

[0108] 4. If Figure 5As shown in the figure, the implementation of the vGPU management module includes:

[0109] (1) The custom scheduler sends the task-node binding information to the Kubelet in the optimal node through the Kubernetes API server;

[0110] (2) The vGPU device plug-in manager in the Kubelet of the optimal node requests the vGPU device plug-in to allocate a vGPU;

[0111] (3) The vGPU device plug-in registers the vGPU and generates a series of device IDs;

[0112] (4) The vGPU device plug-in reports the corresponding vGPU allocation information to Kubelet through the ListAndWatch method, including the vGPU ID;

[0113] (5) Kubelet calls the container runtime to create a Pod, passes resource limits, mounts the corresponding vGPU resources to the Pod and sets environment variables, starts the container, and completes the scheduling;

[0114] (6) At the same time, Kubelet retains the vGPU allocation information in the node cache in the form of annotations and further reports it to the Kubernetes API server;

[0115] (7) The Kubernetes API server stores the received vGPU allocation information in etcd, that is, registers the vGPU using the vGPU ID information in etcd;

[0116] (8) The Kubernetes API server continuously monitors the destruction events and health status of Pods, and the vGPU device plug-in destroys the vGPU corresponding to the Pod that is no longer in the life cycle or is unhealthy to reclaim GPU resources;

[0117] (9) After destroying the vGPU, delete the corresponding information in etcd.

[0118] In the description provided herein, a large number of specific details are described. However, it is understood that embodiments of the present invention can be practiced without these specific details. In some instances, well-known methods, structures, and techniques are not shown in detail so as not to obscure the understanding of this description.

[0119] Although the present invention has been described with respect to a limited number of embodiments, those skilled in the art, having benefit of the foregoing description, will appreciate that other embodiments are contemplated within the scope of the invention thus described. Furthermore, it should be noted that the language used in this specification has been selected primarily for readability and instructional purposes, and not for the purpose of explaining or limiting the subject matter of the present invention.

Claims

1. A method for virtualization isolation and multi-level scheduling of computing resources in a containerized environment, characterized in that: Here are the steps: Step S110: deploy an indicator collector on each working node of the Kubernetes cluster to collect the node computing resource utilization in real time; Step S120: Upon receiving a computing task request submitted by a user, the Kubernetes API server verifies the legitimacy of the fields in the configuration file submitted by the user and stores the result. Step S130: Deploy a custom scheduler on the Kubernetes control plane, execute a multi-level computing resource scheduling algorithm based on the user-submitted configuration file and resource utilization, and find the working node with the highest adaptability as the optimal node; Step S140: The custom scheduler notifies the optimal node to create a vGPU and a Pod, and mounts the vGPU to the corresponding Pod to run the computing task; Step S150: Record the vGPU allocation information of the optimal node, store it, and update the cluster.

2. The method for virtualization isolation and multi-level scheduling of computing resources in a containerized environment according to claim 1 is characterized in that: In step S110, each working node of the Kubernetes cluster includes: Physical GPU resources refer to the physical graphics processor hardware devices installed on the worker nodes, which provide raw computing power support; The vGPU device plug-in is responsible for registering and managing physical GPU resources with Kubelet, and performing vGPU segmentation, status monitoring, and device allocation operations; Virtual GPU resources (vGPUs) are virtualized computing units formed by dividing physical GPUs through vGPU device plug-ins. Each vGPU has independent memory isolation domains and computing core quotas. Pod refers to the smallest task unit scheduled by Kubernetes, which contains one or more containers that share a computing environment; Kubelet, which runs on each working node in the cluster, is responsible for managing the lifecycle of Pods on the running nodes, resource monitoring, node status checking, and container start and stop.

3. The method for virtualization isolation and multi-level scheduling of computing resources in a containerized environment according to claim 1, characterized in that: In step S110, the indicator collector collects node computing resource utilization in real time, including: The indicator collectors include: general computing power indicator collector, intelligent computing power indicator collector and Prometheus service; The general computing power indicator collector is deployed on the node with CPU to collect the utilization of CPU and memory related indicators; The intelligent computing power indicator collector is deployed on the node with GPU, monitoring the static performance indicators such as the number of GPUs, video memory capacity, GPU computing resource utilization, and video memory utilization of the node; The Prometheus service pulls data from the general computing power indicator collector and the intelligent computing power indicator collector every 5 seconds and stores it in the local time series database.

4. The method for virtualization isolation and multi-level scheduling of computing resources in a containerized environment according to claim 1, characterized in that: In step S120, upon receiving a computing task request submitted by a user, the Kubernetes API server verifies the legitimacy of the fields in the configuration file submitted by the user and stores it, including: The user submits a Pod configuration file in YAML format through an API request, which contains a declaration of resource requirements. The Kubernetes API server verifies the legality of the name, type, format, and value range of the fields in the resource requirement declaration based on the custom resource definition. After verification, the Kubernetes API server stores the Pod configuration in etcd.

5. The method for virtualization isolation and multi-level scheduling of computing resources in a containerized environment according to claim 1, characterized in that: In step S130, the Kubernetes control plane includes: The Kubernetes API server, as the only entry point to the Kubernetes API, is the hub for interaction between various components and is responsible for handling all REST requests from users or components within the cluster. Custom scheduler: When a new Pod sends a scheduling request, the custom scheduler will schedule the Pod to the most suitable node based on the resource amount requested by the Pod and the cluster resource utilization according to the custom scheduling algorithm; etcd refers to the Kubernetes backend persistent storage database that stores the status data of the entire cluster.

6. The method for virtualization isolation and multi-level scheduling of computing resources in a containerized environment according to claim 1, characterized in that: In step S130, the multi-level scheduling algorithm for computing resources is implemented as follows: The custom scheduler obtains the Pod configuration file by listening to the Kubernetes API server, parses the resource request (the minimum amount of resources the Pod must guarantee) and the resource limit (the maximum amount of resources the Pod is allowed to use), and obtains the task resource requirements, including the requirements for CPU, memory, number of GPUs, and GPU memory resources. The custom scheduler queries the Prometheus service for resource utilization indicators and receives monitoring data returned by the Prometheus service; At the first level, the custom scheduler filters out unqualified worker nodes from the computing resource cluster based on Pod resource requests and node resource utilization. This includes: when a task requests GPU computing resources, the custom scheduler filters out all worker nodes that do not contain GPUs; when the Pod resource request is greater than the node's remaining computing resources, the custom scheduler filters out this node. At the second level, for nodes that meet the conditions, the scheduling algorithm calculates the fitness of different nodes and tasks to find the optimal node and bind the task; the fitness of the node and the task is composed of the weighted sum of the execution efficiency fitness and the load balancing fitness. The execution efficiency fitness reflects the execution time competitiveness of the task at the node, and the load balancing fitness represents the contribution of task allocation to the global resource load balancing. The weight coefficient is dynamically adjusted according to the real-time status of the task and the node. When the task type is sensitive to execution time, the execution efficiency weight is increased, and when there is significant resource imbalance in the system, the load balancing weight is increased.

7. The method for virtualization isolation and multi-level scheduling of computing resources in a containerized environment according to claim 6, characterized in that: The execution efficiency adaptability and load balancing adaptability are implemented as follows: Execution efficiency fitness: This measure measures the time required for a worker node to execute a task based on the task's CPU or GPU resource requirements. A semi-normalized evaluation model is constructed by comparing the target node's execution time with the longest execution time across the entire cluster. A higher execution efficiency fitness value indicates that the node has better execution efficiency than the cluster as a whole. Load balancing adaptability: This indicator uses the standard deviation of resource utilization across all nodes in the cluster to quantify the degree of global load balancing. Node resource utilization is the weighted sum of CPU and GPU resource utilization, with the weight coefficient adjusted dynamically based on the specific scenario. A larger value for load balancing adaptability indicates that task allocation effectively improves the overall load balancing state of the system.

8. The method for virtualization isolation and multi-level scheduling of computing resources in a containerized environment according to claim 1, characterized in that: In step S140, the optimal node creates a vGPU and a Pod and completes the mounting, including: The custom scheduler sends the task-node binding information to the optimal node; The vGPU device plug-in manager in the Kubelet in the optimal node requests the vGPU device plug-in to allocate a vGPU; The vGPU device plug-in creates a vGPU and reports the corresponding vGPU allocation information to Kubelet, including the vGPU ID; Kubelet creates Pods and mounts them to corresponding vGPUs for scheduling. It also saves vGPU allocation information to the node cache and reports it to the Kubernetes control plane. The Kubernetes API server stores the received vGPU allocation information in etcd.

9. The method for virtualization isolation and multi-level scheduling of computing resources in a containerized environment according to claim 1, characterized in that: Further comprising the steps of: S160, the Kubernetes control plane monitors and obtains Pod destruction events and health status; S170: The vGPU device plug-in destroys the vGPU corresponding to the Pod that is out of lifecycle or unhealthy to reclaim GPU resources.

10. A computing resource virtualization isolation and multi-level scheduling system in a containerized environment, characterized in that: include: Resource monitoring module, task request processing module, custom scheduler module, vGPU module; among them: The resource monitoring module is responsible for real-time monitoring of the resource utilization of the working nodes of the Kubernetes cluster and providing the monitoring data to the custom scheduler module for screening the adaptability of nodes and computing nodes to tasks; The task request processing module is responsible for receiving task requests submitted by users by creating Pod configuration files, verifying the rationality of the configuration file fields and storing them, and providing the task resource requirements to the custom scheduler module for screening the nodes and computing nodes for their suitability with the task; The custom scheduler module first obtains the node's real-time resource utilization indicators from the resource monitoring module. Based on the task resource requirements provided by the task request processing module, it filters out nodes whose current resource availability is less than the task resource requirements. It then uses a custom scheduling algorithm to calculate the execution efficiency and load balancing fitness between each node and task, performing a weighted summation to obtain the overall fitness between the node and task. It then selects the node with the highest overall fitness and binds it to the task. The task-node binding information is then sent to the vGPU management module. After the vGPU management module receives the task-node binding information sent by the custom scheduler, the vGPU device plug-in manager in the Kubelet on the working node applies to the vGPU device plug-in to create a vGPU. The vGPU device plug-in creates a vGPU in the managed GPU resources and returns the ID. The Kubelet creates a Pod corresponding to the task, mounts the Pod on the created vGPU, starts the container to run the computing task, and completes the scheduling of computing resources. At the same time, the vGPU device plug-in reports the created vGPU allocation information to the Kubelet. The Kubelet saves the vGPU allocation information in the Pod cache and further reports it to the control plane for storage. When the corresponding Pod ends or the status is abnormal, the corresponding vGPU is destroyed to achieve resource recovery.

Citation Information

Patent Citations

  • Multi-dimensional resource scheduling method under Kubernetes cluster architecture system

    CN111522639A

  • Container cloud platform GPU resource scheduling method, device and application

    CN115454636A

  • Virtualized GPU (Graphics Processing Unit) scheduling method and device of container system and medium

    CN115904675A

  • GPU task fine-grained scheduling method and related device

    CN116450298A

  • Heterogeneous intelligent computing platform virtualization management system and method

    CN117707693A

Cited By

  • Deep learning training task state monitoring method and device, equipment and medium

    CN121328778A