A Kubernetes cluster, a resource management method thereof, an electronic device, a storage medium and a computer program product
Patent Information
- Application Number
- CN202511587439.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-31
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2045-10-31
AI Technical Summary
[0015]本公开实施例的Kubernetes集群包括:至少一个工作节点,每个工作节点内包括:Kubelet组件、拓扑生成器、Kubernetes设备插件控制器;工作节点内的拓扑生成器,用于生成能够指示工作节点内的第一类型设备和第二类型设备之间的拓扑关系的设备拓扑图;工作节点内的Kubernetes设备插件控制器,用于根据工作节点的设备拓扑图,基于拓扑距离最短原则,确定工作节点内的异构设备对列表,每个异构设备对包括一个第一类型设备和一个第二类型设备;工作节点内的Kubernetes设备插件控制器,用于向工作节点内的Kubelet组件上报工作节点内的异构设备对列表。本公开实施例的Kubernetes集群,能够利用Kubernetes集群中的工作节点内的Kubernetes设备插件控制器,对工作节点内的第一类型设备和第二类型设备,有效实现基于拓扑感知的组合式异构资源管理,进而后续能够实现基于拓扑感知的组合式异构资源绑定分配。
Smart Images

Figure CN121455676B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer technology, and in particular to a Kubernetes cluster and its resource management method, electronic devices, storage media, and computer program products. Background Technology
[0002] Heterogeneous computing refers to integrating various types of computing resources, such as Central Processing Units (CPUs), Graphics Processing Units (GPUs), Field-Programmable Gate Arrays (FPGAs), and Application-Specific Integrated Circuits (ASICs), into the same computing system to improve computing efficiency and resource utilization. Kubernetes, as the core platform for container orchestration, combined with heterogeneous computing, provides powerful resource management and scheduling capabilities for scenarios such as High-Performance Computing (HPC), Artificial Intelligence (AI) training, and edge computing. Therefore, there is an urgent need for a Kubernetes cluster for heterogeneous resource management and scheduling. Summary of the Invention
[0003] In view of this, this disclosure presents a Kubernetes cluster and its resource management method, electronic device, storage medium and computer program product.
[0004] According to one aspect of this disclosure, a Kubernetes cluster is provided, comprising: at least one worker node, each worker node including: a Kubelet component, a topology generator, and a Kubernetes device plugin controller; the topology generator within the worker node is used to generate a device topology graph of the worker node, wherein the device topology graph of the worker node is used to indicate the topological relationship between a first type of device and a second type of device within the worker node; the Kubernetes device plugin controller within the worker node is used to determine a list of heterogeneous device pairs within the worker node based on the device topology graph of the worker node and the principle of shortest topological distance, wherein each heterogeneous device pair includes a first type of device and a second type of device; the Kubernetes device plugin controller within the worker node is used to report the list of heterogeneous device pairs within the worker node to the Kubelet component within the worker node.
[0005] In one possible implementation, each worker node includes: a first type of Kubernetes device plugin and its corresponding first socket file, a second type of Kubernetes device plugin and its corresponding second socket file; a Kubernetes device plugin controller within the worker node, used to communicate with the first type of Kubernetes device plugin within the worker node using the first socket file within the worker node; and a Kubernetes device plugin controller within the worker node, used to communicate with the second type of Kubernetes device plugin within the worker node using the second socket file within the worker node.
[0006] In one possible implementation, a Kubernetes device plugin controller within the worker node is used to obtain a list of first-type devices within the worker node using first-type Kubernetes device plugins within the worker node; a Kubernetes device plugin controller within the worker node is used to obtain a list of second-type devices within the worker node using second-type Kubernetes device plugins within the worker node; and a topology generator within the worker node is used to generate a device topology graph of the worker node based on the first-type device list and the second-type device list within the worker node.
[0007] In one possible implementation, the Kubernetes cluster further includes: a Kubernetes resource scheduler; a Kubelet component within a worker node, used to report a list of heterogeneous device pairs within the worker node to the Kubernetes resource scheduler; the Kubernetes resource scheduler, used to determine a target worker node for the target pod in the Kubernetes cluster based on the resource request of the target pod and the list of heterogeneous device pairs within each worker node, wherein the resource request of the target pod is used to request n heterogeneous device pairs for the target pod, where n is a positive integer greater than or equal to 1; and the Kubernetes resource scheduler, used to bind the target pod to the target worker node.
[0008] In one possible implementation, the Kubelet component within the target worker node is configured to, after determining that the target pod is bound to the target worker node, send the resource request of the target pod to the Kubernetes device plugin controller within the target worker node; the Kubernetes device plugin controller within the target worker node is configured to, based on the resource request of the target pod, select n target heterogeneous device pairs with the shortest topological distance from the list of target heterogeneous device pairs; the Kubernetes device plugin controller within the target worker node is configured to invoke a first type of Kubernetes device plugin and a second type of Kubernetes device plugin within the target worker node to allocate the n target heterogeneous device pairs to the target pod.
[0009] In one possible implementation, the Kubernetes device plugin controller within the target worker node is configured to send the device pair identifier of any target heterogeneous device pair to a first type Kubernetes device plugin and a second type Kubernetes device plugin within the target worker node; the first type Kubernetes device plugin within the target worker node is configured to parse the device pair identifier of the target heterogeneous device pair and allocate the first type device included in the target heterogeneous device pair to the target pod; the second type Kubernetes device plugin within the target worker node is configured to parse the device pair identifier of the target heterogeneous device pair and allocate the second type device included in the target heterogeneous device pair to the target pod.
[0010] In one possible implementation, the first type of device is a GPU device and the second type of device is an RDMA NIC device; or, the first type of device is a GPU device and the second type of device is an FPGA device; or, the first type of device is a GPU device and the second type of device is an NVME device.
[0011] According to another aspect of this disclosure, a resource management method for a Kubernetes cluster is provided. The Kubernetes cluster includes: at least one worker node, each worker node including: a Kubelet component, a topology generator, and a Kubernetes device plugin controller; the method includes: using the topology generator within the worker node to generate a device topology map of the worker node, wherein the device topology map of the worker node is used to indicate the topological relationship between a first type of device and a second type of device within the worker node; based on the device topology map of the worker node, using the Kubernetes device plugin controller within the worker node, determining a list of heterogeneous device pairs within the worker node based on the principle of shortest topological distance, wherein each heterogeneous device pair includes a first type of device and a second type of device; and using the Kubernetes device plugin controller within the worker node to report the list of heterogeneous device pairs within the worker node to the Kubelet component within the worker node.
[0012] According to another aspect of this disclosure, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the above-described method.
[0013] According to another aspect of this disclosure, a non-volatile computer-readable storage medium is provided, on which a computer program is stored, which, when executed by a processor, implements the steps of the above-described method.
[0014] According to another aspect of this disclosure, a computer program product is provided, including a computer program or a non-volatile computer-readable storage medium carrying the computer program, wherein the computer program, when executed by a processor, implements the steps of the above-described method.
[0015] This disclosure's embodiment of the Kubernetes cluster includes: at least one worker node, each worker node including: a Kubelet component, a topology generator, and a Kubernetes device plugin controller; the topology generator within the worker node is used to generate a device topology map that indicates the topological relationship between a first type of device and a second type of device within the worker node; the Kubernetes device plugin controller within the worker node is used to determine a list of heterogeneous device pairs within the worker node based on the device topology map and the principle of shortest topological distance, each heterogeneous device pair including one first type of device and one second type of device; the Kubernetes device plugin controller within the worker node is used to report the list of heterogeneous device pairs within the worker node to the Kubelet component within the worker node. This disclosure's embodiment of the Kubernetes cluster can effectively implement topology-aware, combinatorial heterogeneous resource management for the first type of device and the second type of device within the worker node using the Kubernetes device plugin controller within the worker node, thereby enabling subsequent topology-aware, combinatorial heterogeneous resource binding and allocation.
[0016] Other features and aspects of this disclosure will become clear from the following detailed description of exemplary embodiments with reference to the accompanying drawings. Attached Figure Description
[0017] The accompanying drawings, which are included in and form part of this specification, illustrate exemplary embodiments, features, and aspects of this disclosure together with the specification and serve to explain the principles of this disclosure.
[0018] Figure 1 A block diagram of a Kubernetes cluster according to an embodiment of the present disclosure is shown.
[0019] Figure 2 This diagram illustrates a worker node included in a Kubernetes cluster according to an embodiment of the present disclosure.
[0020] Figure 3 A flowchart illustrating a resource management method for a Kubernetes cluster according to an embodiment of this disclosure is shown.
[0021] Figure 4 A block diagram of an electronic device according to an embodiment of the present disclosure is shown. Detailed Implementation
[0022] Various exemplary embodiments, features, and aspects of this disclosure will now be described in detail with reference to the accompanying drawings. The same reference numerals in the drawings denote elements that have the same or similar functions. Although various aspects of the embodiments are shown in the drawings, they are not necessarily drawn to scale unless specifically indicated otherwise.
[0023] As used herein, the terms “comprising,” “including,” “having,” or variations thereof are open-ended and include one or more of the stated features, integrals, elements, steps, components, or functions, but do not exclude the presence or addition of one or more other features, integrals, elements, steps, components, functions, or groups thereof.
[0024] When an element is referred to as “connected,” “coupled,” “responding,” or a variation thereof relative to another element, it may be directly connected, coupled, or responding to another element, or there may be an intermediate element present.
[0025] Although the terms first, second, third, etc., may be used herein to describe various elements / operations, these elements / operations should not be limited by these terms. These terms are only used to distinguish one element / operation from another. Therefore, without departing from the teachings of the inventive concept, a first element / operation in some embodiments may be referred to as a second element / operation in other embodiments.
[0026] The term “exemplary” as used herein means “serving as an example, embodiment, or illustration.” Any embodiment illustrated herein as “exemplary” is not necessarily to be construed as superior to or better than other embodiments.
[0027] Furthermore, to better illustrate this disclosure, numerous specific details are set forth in the following detailed description. Those skilled in the art will understand that this disclosure can be practiced without certain specific details. In some instances, methods, means, components, and circuits well known to those skilled in the art have not been described in detail in order to highlight the main points of this disclosure.
[0028] Cloud computing is a delivery method that enables the on-demand provision of infrastructure, services, platforms, and applications across networks, rapidly replacing the traditional method of resource sharing via hardwired connections. Cloud-native is a software approach to building, deploying, and managing modern applications in a cloud computing environment. Modern enterprises want to build highly scalable, flexible, and resilient applications that can be rapidly updated to meet customer needs. To do this, they use modern tools and technologies that inherently support application development on cloud infrastructure. These cloud-native technologies enable rapid and frequent application changes without impacting service delivery, thus providing adopters with an innovative competitive advantage.
[0029] A GPU cluster is a computer cluster in which each worker node is equipped with a GPU. Leveraging the computing power of modern GPUs, GPU clusters can perform very fast computations.
[0030] PCIe is a computer bus standard used to connect peripheral devices on a computer motherboard. It allows various devices (such as graphics cards, network interface devices, and storage controllers) to communicate with the motherboard through the same interface. The PCIe bus uses a set of standard interfaces and protocols to achieve communication between devices, providing a flexible way to expand the functionality of a computer. PCIe topology refers to the structural layout of all PCIe devices (such as GPUs and network interface devices) and their connections in a computer system.
[0031] InfiniBand (IB) networking technology is a computer network communication standard for high-performance computing, offering extremely high throughput and extremely low latency for direct data interconnection between computers. Remote Direct Memory Access over Converged Ethernet (RoCE) is an Ethernet-based remote direct memory access protocol that operates on Ethernet networks for efficient data transmission. Employing high-performance network technologies such as IB or RoCE can provide higher bandwidth and lower latency, reduce network congestion, and ensure lossless data transmission even when network congestion occurs.
[0032] Heterogeneous computing refers to integrating various types of computing resources, such as CPUs, GPUs, FPGAs, and ASICs, into the same computing system to improve computing efficiency and resource utilization. In large-scale AI training and inference tasks, heterogeneous computing can fully leverage the advantages of different computing units. For example, GPUs handle matrix operations, CPUs are responsible for scheduling, and Network Interface Cards (NICs) provide high-speed data transmission.
[0033] HPC involves parallel computing and large-scale distributed computing, and is typically used for computationally intensive tasks such as scientific simulations, gene analysis, and weather forecasting. In AI training scenarios, HPC, combined with GPU-accelerated computing and RDMA high-speed interconnects, can significantly improve the training speed of deep learning models.
[0034] Deep learning training relies on large-scale datasets and complex neural networks, typically requiring thousands of GPUs or TPUs for efficient parallel computation. In Kubernetes clusters, proper GPU task scheduling and topology-aware resource allocation can reduce data transfer overhead and accelerate model convergence. AI inference focuses on real-time prediction using trained models, requiring efficient computing resources and low-latency data processing capabilities. Heterogeneous computing resource scheduling (e.g., GPU + FPGA + RDMANIC) can optimize inference performance to meet the needs of applications such as autonomous driving, intelligent recommendation, and real-time speech recognition.
[0035] In large-scale model training and inference scenarios utilizing Kubernetes clusters, using the principle of shortest topological distance as a scheduling strategy for heterogeneous resources can significantly improve computational efficiency and communication performance. Therefore, topology awareness is essential for deep learning training and inference scenarios.
[0036] For example, selecting the GPU device and RDMA NIC device with the shortest topology distance has the following advantages. (1) Reduce data transmission latency between GPU devices and RDMA NIC devices. In distributed training or inference tasks, GPU devices need to transmit data efficiently through RDMA NIC devices (e.g., All Reduce, model parameter synchronization, gradient transfer). Traditional GPU device and RDMA NIC device allocation is random, which may result in cross-PCIE topology allocation, which increases congestion on PCIE and QPI / UPI buses. With a topology-aware scheduling strategy, the GPU device and RDMA NIC device with the shortest topology distance within the same PCIE switch domain can be selected, thereby reducing memory copy overhead and cross-node communication latency. (2) Improve throughput and computational efficiency of AI training tasks. In large-scale AI training (e.g., large models such as GPT-4 and Stable Diffusion), computational tasks often require multiple GPU devices and high-throughput RDMA networks (e.g., InfiniBand, RoCE) to work together. The topology-aware GPU-NIC device binding method can ensure the optimal topology path between the computation flow (GPU computation) and the communication flow (RDMA network communication), making data exchange more efficient and balanced, and improving computation-communication parallelism. Experiments show that reasonable topology binding can reduce the communication overhead of AI training tasks by more than 30%, effectively shortening the training convergence time. (3) Make full use of GPU Direct RDMA network. GPU Direct RDMA technology allows GPU devices to directly access RDMA NIC devices without going through CPU memory, but requires a tight topology between GPU devices and RDMA NIC devices (e.g., GPU devices and RDMA NIC devices are directly connected to a PCIe switch). The topology-aware scheduling strategy binds GPU devices and RDMA NIC devices into GPU-NIC heterogeneous device pairs, which can ensure that GPU devices select the optimal RDMA NIC device, make full use of GPU Direct RDMA, avoid additional data copying, and improve communication bandwidth utilization.
[0037] The Kubernetes device plugin mechanism in a Kubernetes cluster allows custom hardware resources (such as GPUs, RDMA NICs, and TPUs) to be managed and allocated as schedulable resources within the Kubernetes cluster.
[0038] Based on existing Kubernetes device-plugin and topology-aware technologies, the main heterogeneous resource scheduling schemes in the Kubernetes ecosystem are as follows: (1) Individual device plugins. For example, the GPU device plugin only manages GPU devices and does not consider NIC devices; the SR-IOV device plugin only manages NIC devices and does not consider GPU devices. (2) Topology-aware scheduler (NUMA / Topology Manager). The Kubernetes topology-aware scheduler can optimize device allocation based on NUMA nodes. However, since the device plugin itself cannot perform joint scheduling across multiple device plugins, the Kubernetes topology-aware scheduler cannot directly enable the GPU device plugin and the SR-IOV device plugin to work together. Currently, there is no device allocation mode based on the PCIE topology between nodes in the Kubernetes cluster.
[0039] In existing Kubernetes clusters, the Kubernetes resource scheduler (kube-scheduler) can be modified to support affinity scheduling of various heterogeneous devices (e.g., GPU + NIC heterogeneous device combinations). However, the Kubernetes device plugin mechanism mainly operates at the node-level kubelet layer and cannot directly affect the cluster-level Kubernetes resource scheduler's decisions.
[0040] Based on the above existing technologies, the following technical problems exist in current Kubernetes clusters: (1) Different Kubernetes device plugins work independently. For example, GPU device plugins and SR-IOV device plugins run independently and cannot be coordinated for scheduling. (2) Topology affinity is not intelligent enough. The node-level kubelet component simply allocates resources and does not have a heterogeneous resource combination strategy (e.g., GPU-NIC heterogeneous device combination). The problem of cross-PCIE topology binding is serious. However, the topological distance between GPU devices and NIC devices on the PCIE bus directly affects data transmission efficiency. (3) Manual configuration. Currently, users need to manually adjust GPU devices and NIC devices (e.g., node selector), lacking automated scheduling capabilities.
[0041] To address the aforementioned technical issues, this disclosure provides a Kubernetes cluster that effectively enables topology-aware, combinatorial heterogeneous resource management, and subsequently allows for topology-aware, combinatorial heterogeneous resource binding and allocation. The Kubernetes cluster provided in this disclosure is described in detail below.
[0042] Figure 1 A block diagram of a Kubernetes cluster according to an embodiment of this disclosure is shown. Figure 1 As shown, the Kubernetes cluster includes: at least one worker node, and each worker node includes: a Kubelet component, a topology generator, and a Kubernetes device plugin controller; the topology generator within the worker node is used to generate a device topology map of the worker node, wherein the device topology map of the worker node is used to indicate the topological relationship between first-type devices and second-type devices within the worker node; the Kubernetes device plugin controller within the worker node is used to determine a list of heterogeneous device pairs within the worker node based on the device topology map of the worker node and the principle of shortest topological distance, wherein each heterogeneous device pair includes one first-type device and one second-type device; the Kubernetes device plugin controller within the worker node is used to report the list of heterogeneous device pairs within the worker node to the Kubelet component within the worker node.
[0043] The Kubernetes cluster of this disclosure embodiment can effectively implement topology-aware combined heterogeneous resource management for the first type of devices and the second type of devices in the worker nodes of the Kubernetes cluster using the Kubernetes device plugin controller in the worker nodes, and then can subsequently realize topology-aware combined heterogeneous resource binding and allocation.
[0044] In a Kubernetes cluster, worker nodes are allowed to use different hardware devices to achieve heterogeneous computing. For example, worker nodes can deploy heterogeneous hardware devices such as GPUs, RDMA NICs, FPGAs, and NVME devices.
[0045] Within a worker node equipped with a Kubernetes device plugin controller, the Kubernetes device plugin controller can be used to implement combined heterogeneous resource management and scheduling of heterogeneous device pairs within the worker node.
[0046] In one possible implementation, the first type of device is a GPU device and the second type of device is an RDMA NIC device; or, the first type of device is a GPU device and the second type of device is an FPGA device; or, the first type of device is a GPU device and the second type of device is an NVME device.
[0047] Based on the actual heterogeneous resource requirements of the Kubernetes cluster, in the worker node where the Kubernetes device plugin controller is deployed in this embodiment, the heterogeneous resource pair can be a GPU device + RDMA NIC device, a GPU device + FPGA device, a GPU device + NVME device, or other combinations of heterogeneous device pairs. This disclosure does not make any specific limitations on this.
[0048] For any worker node in a Kubernetes cluster, add a Kubernetes device plugin controller within the worker node. The Kubernetes device plugin controller can act as a proxy layer within the worker node, calling other Kubernetes device plugins within the worker node that manage specific hardware devices.
[0049] For any worker node in a Kubernetes cluster, the Kubernetes device plugin mechanism enables each Kubernetes device plugin to create a corresponding socket file in the / var / lib / kubelet / device-plugins directory within the worker node when it starts, so as to report the device type of the Kubernetes device plugin to the Kubelet component within the worker node.
[0050] In one possible implementation, each worker node includes: a first type of Kubernetes device plugin and its corresponding first socket file, a second type of Kubernetes device plugin and its corresponding second socket file; a Kubernetes device plugin controller within the worker node, used to communicate with the first type of Kubernetes device plugin within the worker node using the first socket file within the worker node; and a Kubernetes device plugin controller within the worker node, used to communicate with the second type of Kubernetes device plugin within the worker node using the second socket file within the worker node.
[0051] Since each worker node in a Kubernetes cluster has both Type 1 and Type 2 devices deployed, in order to manage resources on each worker node, it is necessary to deploy a Type 1 Kubernetes device plugin for managing Type 1 devices and a Type 2 Kubernetes device plugin for managing Type 2 devices on each worker node.
[0052] like Figure 1As shown, the worker node has a first type of Kubernetes device plugin and a second type of Kubernetes device plugin deployed within it.
[0053] For any worker node in a Kubernetes cluster, when the first type of Kubernetes device plugin and the second type of Kubernetes device plugin are started, the first socket file corresponding to the first type of Kubernetes device plugin and the second socket file corresponding to the second type of Kubernetes device plugin will be created in the / var / lib / kubelet / device-plugins directory within the worker node.
[0054] For any worker node in a Kubernetes cluster, the Kubernetes device plugin controller within the worker node can determine the number of Kubernetes device plugins within the worker node by obtaining the number of socket files in the / var / lib / kubelet / device-plugins directory within the worker node.
[0055] For any worker node in a Kubernetes cluster, the Kubernetes device plugin controller within the worker node can communicate with the Kubernetes device plugins based on the socket file of each Kubernetes device plugin within the worker node.
[0056] For any worker node in a Kubernetes cluster, the worker node includes: a first type of Kubernetes device plugin and its corresponding first socket file, and a second type of Kubernetes device plugin and its corresponding second socket file. In this case, the Kubernetes device plugin controller within the worker node is used to communicate with the first type of Kubernetes device plugin within the worker node using the first socket file; and to communicate with the second type of Kubernetes device plugin within the worker node using the second socket file.
[0057] Within any worker node in a Kubernetes cluster, the Kubernetes Device Plugin Controller acts as a device plugin proxy layer within the worker node. It serves as an intermediary between the Kubelet component within the worker node and the device plugins corresponding to different underlying hardware devices. It can shield the differences in underlying hardware, provide a unified heterogeneous resource management interface to the Kubelet component, and support dynamic registration and resource management of device plugins corresponding to different hardware devices.
[0058] For example, the first type of device is a GPU device, and the first type of Kubernetes device plugin is a GPU device plugin; the second type of device is an RDMA NIC device, and the second type of Kubernetes device plugin is an SR-IOV device plugin. The Kubernetes device plugin controller communicates with the GPU device plugin using the socket file of the GPU device plugin; and communicates with the SR-IOV device plugin using the socket file of the SR-IOV device plugin.
[0059] Within any worker node in a Kubernetes cluster, the Kubernetes Device Plugin Controller acts as a device plugin proxy layer within the worker node. It serves as an intermediary between the Kubelet component within the worker node and the underlying GPU device plugins and SR-IOV device plugins. It can shield the differences in the underlying hardware, provide a unified heterogeneous resource management interface to the Kubelet component, and support dynamic registration and resource management of GPU device plugins and SR-IOV device plugins.
[0060] In one possible implementation, a Kubernetes device plugin controller within the worker node is used to obtain a list of first-type devices within the worker node using first-type Kubernetes device plugins within the worker node; a Kubernetes device plugin controller within the worker node is used to obtain a list of second-type devices within the worker node using second-type Kubernetes device plugins within the worker node; and a topology generator within the worker node is used to generate a device topology graph of the worker node based on the first-type device list and the second-type device list within the target worker node.
[0061] For any worker node in the Kubernetes cluster, the Kubernetes device plugin controller within the worker node calls the first type of Kubernetes device plugin within the worker node to obtain a list of first-type devices; and calls the second type of Kubernetes device plugin within the worker node to obtain a list of second-type devices. Then, using the topology generator within the worker node, a device topology graph is generated based on the first and second type device lists to indicate the topological relationships between the first and second type devices within the worker node.
[0062] For any worker node in a Kubernetes cluster, after determining the device topology of the worker node, the first type of devices and the second type of devices within the worker node are paired based on the principle of shortest topological distance, resulting in a list of heterogeneous device pairs including multiple first type-second type device pairs. Then, the list of heterogeneous device pairs within the worker node is reported to the Kubelet component within the worker node, effectively realizing topology-aware combined heterogeneous resource management of first type-second type device pairs within the worker node.
[0063] Within a Kubernetes cluster, on worker nodes with a Kubernetes device plugin controller deployed, only Type 1 and Type 2 device pairs can be used; Type 1 and Type 2 devices cannot be used independently. However, on other worker nodes within the Kubernetes cluster that do not have a Kubernetes device plugin controller deployed, Type 1 and Type 2 devices can be used independently.
[0064] Compared to the existing technology that deploys the Kubernetes device plugin controller in a separate worker node in a Kubernetes cluster, the present disclosure embodiment deploys the Kubernetes device plugin controller in each worker node, eliminating the need to consider data synchronization between the Kubernetes device plugin controller in the worker node and the device plugins corresponding to different types of hardware devices.
[0065] Taking the first type of device as a GPU device and the second type of device as an RDMA NIC device as an example: For any worker node in the Kubernetes cluster, the Kubernetes device plugin controller within the worker node calls the GPU device plugin within the worker node to obtain the list of GPU devices within the worker node. The Kubernetes device plugin controller within the worker node then calls the SR-IOV device plugin within the worker node to obtain the list of RDMA NIC devices within the worker node.
[0066] Figure 2 This diagram illustrates a worker node included in a Kubernetes cluster according to an embodiment of the present disclosure. Figure 2 As shown, the Kubernetes device plugin controller calls the list and watch interface of the GPU device plugin to obtain the list of GPU devices in the worker node, and calls the list and watch interface of the SR-IOV device plugin to obtain the list of RDMA NIC devices in the worker node.
[0067] Then, the Kubernetes device plugin controller in the worker node calls the topology generator in the worker node to generate a GPU-RDMA NIC device topology map of the worker node based on the list of GPU devices and the list of RDMA NIC devices in the worker node. The GPU-RDMA NIC device topology map of the worker node is used to indicate the topological relationship between each GPU device and each RDMA NIC device in the worker node.
[0068] like Figure 2 As shown, the Kubernetes device plugin controller calls the topology generator to generate a GPU-RDMA NIC device topology diagram for the worker nodes.
[0069] After determining the GPU-RDMA NIC device topology of the worker node, the Kubernetes device plugin controller within the worker node pairs the GPU devices and RDMA NIC devices based on the principle of shortest topological distance, resulting in a list of GPU-RDMA NIC device pairs. Each GPU device in this list is matched with the optimal RDMA NIC device that has the shortest topological distance.
[0070] Within a Kubernetes cluster, on worker nodes with a Kubernetes Device Controller deployed, only GPU-RDMA NIC device pairs can be used; GPU devices and RDMA NIC devices cannot be used independently. However, on other worker nodes within the Kubernetes cluster without a Kubernetes Device Controller deployed, GPU devices and RDMA NIC devices can be used independently.
[0071] Compared to the existing technology that deploys the Kubernetes device plugin controller in a separate worker node in a Kubernetes cluster, the method of deploying the Kubernetes device plugin controller in each worker node in this disclosure embodiment eliminates the need to consider data synchronization between the Kubernetes device plugin controller in the worker node, the GPU device plugin corresponding to the GPU device, and the SR-IOV device plugin corresponding to the RDMANIC device.
[0072] In one possible implementation, the Kubernetes device plugin controller is a standard Kubernetes device plugin.
[0073] Since the Kubernetes Device Plugin Controller is also a standard Kubernetes device plugin, it also has its corresponding socket file.
[0074] For any worker node in a Kubernetes cluster, the Kubelet component within the worker node can communicate with the Kubernetes device plugin controller using the socket file corresponding to the Kubernetes device plugin controller within the worker node.
[0075] Taking the first type of device as a GPU device and the second type of device as an RDMA NIC device as an example, for any worker node in the Kubernetes cluster, after the Kubernetes device plugin controller within the worker node determines the list of GPU-RDMA NIC device pairs within the worker node, the Kubernetes device plugin controller within the worker node starts the gRPC service based on its corresponding socket file, and reports the list of GPU-RDMA NIC device pairs within the worker node to the Kubelet component within the worker node based on the gRPC service.
[0076] like Figure 2 As shown, the Kubernetes device plugin controller uses the list and watch interface to report the list of GPU-RDMA NIC device pairs within the worker nodes to the Kubelet component.
[0077] In one possible implementation, the Kubernetes cluster also includes: a Kubernetes resource scheduler; a Kubelet component within the worker nodes, used to report a list of heterogeneous device pairs within the worker nodes to the Kubernetes resource scheduler; the Kubernetes resource scheduler, used to determine a target worker node for the target pod in the Kubernetes cluster based on the resource request of the target pod and the list of heterogeneous device pairs within each worker node, wherein the resource request of the target pod is used to request n heterogeneous device pairs for the target pod, where n is a positive integer greater than or equal to 1; and the Kubernetes resource scheduler, used to bind the target pod to the target worker node.
[0078] like Figure 1 As shown, a Kubernetes cluster also includes the Kubernetes resource scheduler (kube-scheduler), which is used for cluster-level resource scheduling within the Kubernetes cluster.
[0079] The Kubelet component within each worker node in a Kubernetes cluster reports a list of heterogeneous device pairs within the worker node to the Kubernetes resource scheduler, enabling the Kubernetes resource scheduler to flexibly schedule heterogeneous device pairs in the Kubernetes cluster.
[0080] When a user needs to request n heterogeneous device pairs (a combination of Type I and Type II devices) for a target pod in a Kubernetes cluster, a resource request for the target pod is generated. Then, the Kubernetes resource scheduler in the Kubernetes cluster determines a target worker node for the target pod based on the list of heterogeneous device pairs within each worker node in the Kubernetes cluster. The target worker node's list of heterogeneous device pairs includes more than n heterogeneous device pairs, meaning the target worker node can satisfy the target pod's resource request. Finally, the Kubernetes resource scheduler binds the target pod to the target worker node.
[0081] In one possible implementation, the Kubelet component within the target worker node, after determining that the target pod is bound to the target worker node, sends the resource request of the target pod to the Kubernetes device plugin controller within the target worker node; the Kubernetes device plugin controller within the target worker node, based on the resource request of the target pod, selects n target heterogeneous device pairs with the shortest topological distance from the list of target heterogeneous device pairs; the Kubernetes device plugin controller within the target worker node, invoking the first type of Kubernetes device plugin and the second type of Kubernetes device plugin within the target worker node, allocates n target heterogeneous device pairs to the target pod.
[0082] After the target pod is bound to the target worker node, the following resource allocation operation is performed within the target worker node: The Kubelet component sends the resource request of the target pod to the Kubernetes device plugin controller through the socket file of the Kubernetes device plugin controller; then, the Kubernetes device plugin controller selects the n target heterogeneous device pairs with the shortest topological distance from the list of heterogeneous device pairs within the target worker node according to the resource request of the target pod, and allocates the first type device and the second type device from the n target heterogeneous device pairs to the target pod.
[0083] In one possible implementation, the Kubernetes device plugin controller within the target worker node is used to send the device pair identifier of any target heterogeneous device pair to the first type Kubernetes device plugin and the second type Kubernetes device plugin within the target worker node; the first type Kubernetes device plugin within the target worker node is used to parse the device pair identifier of the target heterogeneous device pair and assign the first type device included in the target heterogeneous device pair to the target pod; the second type Kubernetes device plugin within the target worker node is used to parse the device pair identifier of the target heterogeneous device pair and assign the second type device included in the target heterogeneous device pair to the target pod.
[0084] Within the target worker node, the Kubernetes device plugin controller acts as a proxy layer. For any one of the n target heterogeneous device pairs determined for the target pod, it sends the device pair identifier of the target heterogeneous device pair to the first type of Kubernetes device plugin and the second type of Kubernetes device plugin.
[0085] The first type of Kubernetes device plugin parses the device pair identifier of the target heterogeneous device pair, assigns the first type of devices included in the target heterogeneous device pair to the target pod, and returns allocation completion information to the Kubernetes device plugin controller; and the second type of Kubernetes device plugin is used to parse the device pair identifier of the target heterogeneous device pair, assign the second type of devices included in the target heterogeneous device pair to the target pod, and return allocation completion information to the Kubernetes device plugin controller.
[0086] After receiving the allocation completion information of the first type of device in each target heterogeneous device pair returned by the first type of Kubernetes device plugin, and the allocation completion information of the second type of device in each target heterogeneous device pair returned by the second type of Kubernetes device plugin, the Kubernetes device plugin controller returns the resource allocation completion information of the target pod to the Kubelet component. At this point, the resource allocation of the heterogeneous device pair for the target pod within the target worker node is completed.
[0087] Taking the first type of device as a GPU device and the second type of device as an RDMA NIC device as an example, when a user needs to request n GPU-RDMA NIC device pairs for a target pod in the Kubernetes cluster, a resource request for the target pod is generated. Then, the Kubernetes resource scheduler in the Kubernetes cluster determines the target worker node for the target pod based on the list of GPU-RDMA NIC device pairs within each worker node in the Kubernetes cluster. The target worker node's list of GPU-RDMA NIC device pairs includes more than n pairs, meaning the target worker node can satisfy the target pod's resource request. Subsequently, the Kubernetes resource scheduler binds the target pod to the target worker node.
[0088] After binding the target pod to the target worker node, the following resource allocation operations are performed within the target worker node: The Kubelet component sends the target pod's resource request to the Kubernetes device plugin controller through the socket file of the Kubernetes device plugin controller; then, based on the target pod's resource request, the Kubernetes device plugin controller selects the n target GPU-RDMANIC device pairs with the shortest topology distance from the list of GPU-RDMA NIC device pairs within the target worker node, and allocates the GPU device and RDMA NIC device from the n target GPU-RDMA NIC device pairs to the target pod.
[0089] like Figure 2 As shown, the Kubelet component calls the allocate interface of the Kubernetes device plugin controller to send the resource request of the target pod to the Kubernetes device plugin controller.
[0090] Within the target worker node, the Kubernetes device plugin controller acts as a proxy layer. For any one of the n target GPU-RDMA NIC device pairs determined for the target pod, it sends the device pair identifier of the target GPU-RDMA NIC device pair to the GPU device plugin and the SR-IOV device plugin.
[0091] like Figure 2As shown, the Kubernetes device plugin controller calls the allocate interface of the GPU device plugin to send the device pair identifier of the target GPU-RDMA NIC device pair to the GPU device plugin, and calls the allocate interface of the SR-IOV device plugin to send the device pair identifier of the target GPU-RDMA NIC device pair to the SR-IOV device plugin.
[0092] The GPU device plugin parses the device pair identifier of the target GPU-RDMA NIC device pair, assigns the GPU devices included in the target GPU-RDMA NIC device pair to the target pod, and returns allocation completion information to the Kubernetes device plugin controller; and the SR-IOV device plugin parses the device pair identifier of the target GPU-RDMA NIC device pair, assigns the RDMA NIC devices included in the target GPU-RDMA NIC device pair to the target pod, and returns allocation completion information to the Kubernetes device plugin controller.
[0093] After receiving the allocation completion information of the GPU devices in each target GPU-RDMA NIC device pair returned by the GPU device plugin and the allocation completion information of the RDMA NIC devices in each target GPU-RDMA NIC device pair returned by the SR-IOV device plugin, the Kubernetes device plugin controller returns the resource allocation completion information of the target pod to the Kubelet component. At this point, the resource allocation of heterogeneous device pairs for the target pod within the target worker node is completed.
[0094] Assigning target GPU-RDMA NIC devices with the shortest distance to target pods based on topology awareness can effectively improve data transmission efficiency and enhance performance in large-scale training.
[0095] It is understood that the various embodiments mentioned above in this disclosure can be combined with each other to form combined embodiments without violating the principle and logic. Due to space limitations, this disclosure will not elaborate further. Those skilled in the art will understand that in the above embodiments of specific implementation methods, the specific execution order of each step should be determined by its function and possible internal logic.
[0096] In addition, this disclosure also provides resource management methods, electronic devices, computer-readable storage media, and programs for Kubernetes clusters.
[0097] Figure 3A flowchart illustrating a resource management method for a Kubernetes cluster according to an embodiment of this disclosure is shown. The Kubernetes cluster includes: at least one worker node, each worker node including: a Kubelet component, a topology generator, and a Kubernetes device plugin controller. Figure 3 As shown, the method may include:
[0098] In step S31, a device topology diagram of the working node is generated using the topology generator within the working node. The device topology diagram of the working node is used to indicate the topological relationship between the first type of devices and the second type of devices within the working node.
[0099] In step S32, based on the device topology diagram of the worker node, the Kubernetes device plugin controller within the worker node is used to determine a list of heterogeneous device pairs within the worker node based on the principle of shortest topological distance. Each heterogeneous device pair includes a first type of device and a second type of device.
[0100] In step S33, the Kubernetes device plugin controller within the worker node is used to report the list of heterogeneous device pairs within the worker node to the Kubelet component within the worker node.
[0101] In one possible implementation, each worker node includes: a first type of Kubernetes device plugin and its corresponding first socket file, and a second type of Kubernetes device plugin and its corresponding second socket file;
[0102] The method also includes:
[0103] The Kubernetes device plugin controller within the worker node communicates with the first type of Kubernetes device plugin within the worker node using the first socket file within the worker node.
[0104] The Kubernetes device plugin controller within the worker node communicates with the second type of Kubernetes device plugin within the worker node using the second socket file within the worker node.
[0105] In one possible implementation, a topology generator within the worker node is used to generate the device topology diagram of the worker node, including:
[0106] The Kubernetes device plugin controller within the worker node uses the first type of Kubernetes device plugin within the worker node to obtain the list of the first type of devices within the worker node;
[0107] The Kubernetes device plugin controller within the worker node uses the second type of Kubernetes device plugin within the worker node to obtain a list of second type devices within the worker node;
[0108] Using the topology generator within the worker node, generate the device topology diagram of the worker node based on the first type of device list and the second type of device list within the worker node.
[0109] In one possible implementation, the Kubernetes cluster also includes: a Kubernetes resource scheduler;
[0110] The method also includes:
[0111] Use the Kubelet component within the worker node to report the list of heterogeneous device pairs within the worker node to the Kubernetes resource scheduler;
[0112] Using the Kubernetes resource scheduler, the target worker node is determined for the target pod in the Kubernetes cluster based on the resource request of the target pod and the list of heterogeneous device pairs in each worker node. The resource request of the target pod is used to apply for n heterogeneous device pairs for the target pod, where n is a positive integer greater than or equal to 1.
[0113] Use the Kubernetes resource scheduler to bind the target pod to the target worker node.
[0114] In one possible implementation, the method further includes:
[0115] After determining that the target pod is bound to the target worker node, the Kubelet component within the target worker node is used to send the resource requests of the target pod to the Kubernetes device plugin controller within the target worker node;
[0116] Using the Kubernetes device plugin controller within the target worker node, select the n target heterogeneous device pairs with the shortest topological distance from the target pod's resource request list;
[0117] By utilizing the Kubernetes device plugin controller within the target worker node, the first type of Kubernetes device plugin and the second type of Kubernetes device plugin within the target worker node are invoked to allocate the n target heterogeneous device pairs to the target pod.
[0118] In one possible implementation, the method further includes:
[0119] For any target heterogeneous device pair, the device pair identifier of the target heterogeneous device pair is sent to the first type of Kubernetes device plugin and the second type of Kubernetes device plugin in the target worker node using the Kubernetes device plugin controller in the target worker node;
[0120] Using the first type of Kubernetes device plugin within the target worker node, the device pair identifier of the target heterogeneous device pair is parsed, and the first type of device included in the target heterogeneous device pair is assigned to the target pod;
[0121] Using the second type of Kubernetes device plugin within the target worker node, the device pair identifier of the target heterogeneous device pair is parsed, and the second type of device included in the target heterogeneous device pair is assigned to the target pod.
[0122] In one possible implementation,
[0123] The first type of device is a GPU device, and the second type of device is an RDMA NIC device; or,
[0124] The first type of device is a GPU device, and the second type of device is an FPGA device; or,
[0125] The first type of device is the GPU device, and the second type of device is the NVME device.
[0126] In some embodiments, the functions or modules of the apparatus provided in this disclosure can be used to perform the methods described in the above method embodiments. The specific implementation can be referred to the description of the above method embodiments, and for the sake of brevity, it will not be repeated here.
[0127] This disclosure also provides an electronic device, including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the above method.
[0128] This disclosure also provides a non-volatile computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the above-described method.
[0129] This disclosure also provides a computer program product, including a computer program or a non-volatile computer-readable storage medium carrying the computer program, wherein the computer program, when executed by a processor, implements the steps of the above method.
[0130] Figure 4 A block diagram of an electronic device according to an embodiment of the present disclosure is shown. (Refer to...) Figure 4 Device 1900 can be provided as a server or terminal device. (See reference...) Figure 4 The apparatus 1900 includes a processing component 1922, which further includes one or more processors, and memory resources represented by memory 1932 for storing instructions, such as application programs, that can be executed by the processing component 1922. The application programs stored in memory 1932 may include one or more modules, each corresponding to a set of instructions. Furthermore, the processing component 1922 is configured to execute instructions to perform the methods described above.
[0131] Device 1900 may also include a power supply component 1926 configured to perform power management of device 1900, a wired or wireless network interface 1950 configured to connect device 1900 to a network, and an input / output interface 1958 (I / O interface). Device 1900 can operate on an operating system, such as Windows Server, stored in memory 1932. TM macOS X TM Unix TM Linux TM FreeBSD TM Or similar.
[0132] In an exemplary embodiment, a non-volatile computer-readable storage medium is also provided, such as a memory 1932 including computer program instructions that can be executed by a processing component 1922 of the device 1900 to perform the above-described method.
[0133] Computer-readable storage media can be tangible devices capable of holding and storing programs / instructions used by instruction execution devices. Computer-readable storage media can be, for example—but not limited to—electrical storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any suitable combination of the foregoing. The computer-readable storage media used herein are not to be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.
[0134] The computer program (or computer-readable program instructions) described herein can be downloaded from a computer-readable storage medium to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage medium in the respective computing / processing device.
[0135] The computer program (or computer program instructions) used to perform the operations of this disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry, such as programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), is personalized by utilizing state information from the computer-readable program instructions to implement various aspects of this disclosure.
[0136] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0137] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processor of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0138] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0139] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those shown in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0140] The various embodiments of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical application, or technical improvements to the embodiments in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.
Claims
1. A Kubernetes cluster, characterized in that, include: At least one worker node, each worker node includes: a Kubelet component, a topology generator, and a Kubernetes device plugin controller; A topology generator within a worker node is used to generate a device topology diagram of the worker node, wherein the device topology diagram of the worker node is used to indicate the topological relationship between the first type of devices and the second type of devices within the worker node; The Kubernetes device plugin controller within the worker node is used to determine a list of heterogeneous device pairs within the worker node based on the device topology graph of the worker node and the principle of shortest topological distance. Each heterogeneous device pair includes one first-type device and one second-type device. The Kubernetes device plugin controller within the worker node is used to report the list of heterogeneous device pairs within the worker node to the Kubelet component within the worker node.
2. The Kubernetes cluster according to claim 1, characterized in that, Each worker node includes: a first type of Kubernetes device plugin and its corresponding first socket file, and a second type of Kubernetes device plugin and its corresponding second socket file; The Kubernetes device plugin controller within the worker node is used to communicate with the first type of Kubernetes device plugin within the worker node using the first socket file within the worker node; The Kubernetes device plugin controller within the worker node is used to communicate with the second type of Kubernetes device plugin within the worker node using a second socket file within the worker node.
3. The Kubernetes cluster according to claim 2, characterized in that, The Kubernetes device plugin controller within the worker node is used to obtain a list of first-type devices within the worker node using the first-type Kubernetes device plugin within the worker node; The Kubernetes device plugin controller within the worker node is used to obtain a list of second-type devices within the worker node using second-type Kubernetes device plugins within the worker node; The topology generator within the worker node is used to generate a device topology diagram of the worker node based on the first type of device list and the second type of device list within the worker node.
4. The Kubernetes cluster according to claim 2, characterized in that, The Kubernetes cluster also includes: a Kubernetes resource scheduler; The Kubelet component within the worker node is used to report the list of heterogeneous device pairs within the worker node to the Kubernetes resource scheduler; The Kubernetes resource scheduler is used to determine a target worker node for the target pod in the Kubernetes cluster based on the resource request of the target pod and the list of heterogeneous device pairs in each worker node. The resource request of the target pod is used to apply for n heterogeneous device pairs for the target pod, where n is a positive integer greater than or equal to 1. The Kubernetes resource scheduler is used to bind the target pod to the target worker node.
5. The Kubernetes cluster according to claim 4, characterized in that, The Kubelet component within the target worker node is used to send the resource request of the target pod to the Kubernetes device plugin controller within the target worker node after determining that the target pod is bound to the target worker node. The Kubernetes device plugin controller within the target worker node is used to select the n target heterogeneous device pairs with the shortest topological distance from the target heterogeneous device pair list based on the resource requests of the target pod. The Kubernetes device plugin controller within the target worker node is used to call the first type of Kubernetes device plugin and the second type of Kubernetes device plugin within the target worker node to allocate the n target heterogeneous device pairs to the target pod.
6. The Kubernetes cluster according to claim 5, characterized in that, The Kubernetes device plugin controller within the target worker node is used to send the device pair identifier of any target heterogeneous device pair to the first type of Kubernetes device plugin and the second type of Kubernetes device plugin within the target worker node; The first type of Kubernetes device plugin within the target worker node is used to parse the device pair identifier of the target heterogeneous device pair and assign the first type of device included in the target heterogeneous device pair to the target pod; The second type of Kubernetes device plugin within the target worker node is used to parse the device pair identifier of the target heterogeneous device pair and assign the second type of device included in the target heterogeneous device pair to the target pod.
7. The Kubernetes cluster according to any one of claims 1 to 6, characterized in that, The first type of device is a GPU device, and the second type of device is an RDMA NIC device; or, The first type of device is a GPU device, and the second type of device is an FPGA device; or, The first type of device is the GPU device, and the second type of device is the NVME device.
8. A resource management method for a Kubernetes cluster, characterized in that, The Kubernetes cluster includes: at least one worker node, and each worker node includes: a Kubelet component, a topology generator, and a Kubernetes device plugin controller; the method includes: Using the topology generator within the worker node, a device topology diagram of the worker node is generated, wherein the device topology diagram of the worker node is used to indicate the topological relationship between the first type of devices and the second type of devices within the worker node; Based on the device topology of the worker node, the list of heterogeneous device pairs within the worker node is determined using the Kubernetes device plugin controller within the worker node, based on the principle of shortest topological distance. Each heterogeneous device pair includes one first-type device and one second-type device. The Kubernetes device plugin controller within the worker node is used to report the list of heterogeneous device pairs within the worker node to the Kubelet component within the worker node.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement the steps of the method of claim 8.
10. A non-volatile computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method of claim 8.
11. A computer program product comprising a computer program, or a non-volatile computer-readable storage medium carrying a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method of claim 8.
Citation Information
Patent Citations
Resource scheduling method and device, equipment and storage medium
CN113377520A
Topology-aware provisioning of hardware accelerator resources in a distributed environment
US20190312772A1