Non-intrusive resource allocation method and device based on Kubernetes cluster and storage medium
By creating a proxy server in the Kubernetes cluster node to segment GPU resources, the problem of inefficient resource allocation under traditional resource management methods is solved, and refined management and efficient utilization of GPU resources are achieved.
Patent Information
- Application Number
- CN202510437547.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2045-04-09
AI Technical Summary
The traditional resource management method based on nodes cannot effectively distinguish the performance differences of heterogeneous devices, resulting in inefficient resource allocation, inability to meet the needs of different types of tasks, and inability to avoid resource contention in a multi-tenant environment, resulting in low GPU resource utilization.
In the Kubernetes cluster node, after receiving the registration request for the GPU resource device plug-in, a proxy server for the proxy GPU resource device plug-in is created based on the preset configuration information, and the GPU resources on the node are divided through the proxy server to achieve refined management of GPU resources.
The GPU resources are segmented through the proxy server, which realizes refined management of GPU resources on the node, improves resource utilization, avoids resource competition, and meets the needs of different types of tasks.
Smart Images

Figure CN119961006A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of cloud-native scheduling, and in particular to a non-intrusive resource allocation method, device, and storage medium based on a Kubernetes cluster. Background Art
[0002] Kubernetes (often referred to as k8s) cluster is an open source platform for automating the deployment, expansion and management of containerized applications. With the rapid development of cloud native technologies represented by k8s, more and more deep learning and high-performance computing tasks are beginning to be uniformly orchestrated and managed through the k8s platform. In addition, with the advancement of computing hardware, GPUs, as core computing resources, are widely used in modern deep learning and high-performance computing scenarios. However, due to its high cost and the large number of suppliers, in order to maximize resource utilization and reduce costs, many k8s clusters adopt heterogeneous GPU architectures, that is, multiple models or manufacturers of GPU devices exist in the cluster at the same time. In actual deployment, k8s allows users to dynamically deploy, manage and expand applications in the cluster. GPU vendors follow the plug-in extension mechanism standard of k8s and provide their own device plug-ins. Through their own device plug-ins, GPU resources are registered in the k8s cluster and can be scheduled and allocated. In a Kubernetes (K8s) cluster, a node is the basic unit of the cluster, which is used to run and manage containerized applications (Pods). K8s manages and schedules GPU resources through a plug-in extension mechanism. It usually deploys the same type of GPU on a single node, that is, only GPU devices of the same model or manufacturer are deployed on a node. Each node is considered to be the basic unit of the resource pool, and the required GPU devices are directly called on each node through the plug-in extension mechanism to implement resource calls.
[0003] However, within the node, the types and performance differences of cloud computing devices are increasing. Especially with the popularization of heterogeneous devices such as GPU, TPU, FPGA, etc., different types of devices (CPU, GPU, TPU, FPGA, etc.) have different performance and uses. The traditional resource management method based on nodes cannot effectively distinguish these differences, resulting in inefficient resource allocation. Secondly, applications and tasks have different requirements for resources. Some tasks may rely more on the computing power of the CPU, while others may require the parallel processing capabilities of the GPU. The traditional resource management method based on nodes cannot meet the needs of different types of tasks. In addition, in a multi-tenant environment, it is necessary to ensure resource isolation between different tenants to avoid resource contention and performance degradation. The traditional resource management method based on nodes cannot avoid resource contention, resulting in low GPU resource utilization.
[0004] There is no effective solution to the problem of low resource utilization in related technologies. Summary of the invention
[0005] In this embodiment, a resource allocation method, device, electronic device and storage medium are provided to solve the problem of low resource utilization in related technologies.
[0006] In a first aspect, a non-intrusive resource allocation method based on a Kubernetes cluster is provided in this embodiment, including: In the K8s cluster node, after receiving the registration request of the GPU resource device plug-in, a proxy server for the GPU resource device plug-in is created according to the preset configuration information; The GPU resources on the node are divided by the proxy server, wherein the GPU resource status on the node is obtained by the deployed proxy service.
[0007] In some of the embodiments, before receiving the registration request of the GPU resource device plug-in, the method further includes: Initialize the configuration information of the GPU resources on the node.
[0008] In some of the embodiments, after the GPU resources on the node are segmented by the proxy server, the method further includes: The corresponding configuration strategy is executed according to the preset configuration information through the interface extension point mechanism, wherein each configuration strategy corresponds to a configuration item function in the preset configuration information.
[0009] In some of the embodiments, the configuration item function includes grouping information of the GPU resources and hidden switch configuration of the GPU resources.
[0010] In some embodiments, creating a proxy server that acts as a proxy for the GPU resource device plug-in according to preset configuration information includes: Reading the preset configuration information, and determining the number of groups of the GPU resources according to the configuration information; Create a number of proxy servers that are proxies for the GPU resource device plug-ins that is the same as the number of groups of the GPU resources.
[0011] In some of the embodiments, after creating the same number of proxy servers as the number of groups of the GPU resources that proxy the GPU resource device plug-ins, the method further includes: The proxy server registers the resource specification name with the main node agent on the node to obtain a new resource specification name, so that the k8s cluster can retrieve GPU resources according to the new resource specification name.
[0012] In some of the embodiments, when the GPU resource device plug-in on the node is offline, the proxy server corresponding to the GPU resource device plug-in is destroyed.
[0013] In a second aspect, in this embodiment, a non-intrusive resource allocation device based on a Kubernetes cluster is provided, including: a creation module and a segmentation module, wherein: The creation module is used to create a proxy server that acts as a proxy for the GPU resource device plug-in in the K8s cluster node according to preset configuration information after receiving a registration request for the GPU resource device plug-in; The segmentation module is used to segment the GPU resources on the node through the proxy server, wherein the GPU resources on the node are obtained through the deployed proxy service.
[0014] In a third aspect, an electronic device is provided in this embodiment, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the non-intrusive resource allocation method based on the Kubernetes cluster described in the first aspect is implemented.
[0015] In a fourth aspect, a storage medium is provided in this embodiment, on which a computer program is stored, and when the program is executed by a processor, the non-intrusive resource allocation method based on the Kubernetes cluster described in the first aspect above is implemented.
[0016] Compared with the related art, the non-intrusive resource allocation method based on Kubernetes cluster provided in this embodiment, after receiving the registration request of the GPU resource device plug-in in the K8s cluster node, creates a proxy server for the proxy GPU resource device plug-in according to preset configuration information; the GPU resources on the node are divided by the proxy server, wherein the GPU resources on the node are obtained through the deployed proxy service, thereby realizing the refined management of GPU resources on the node and improving the utilization rate of resources.
[0017] Details of one or more embodiments of the present application are set forth in the following drawings and description to make other features, objects, and advantages of the present application more readily apparent. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings: Figure 1 This is a hardware structure block diagram of a terminal of the non-intrusive resource allocation method based on a Kubernetes cluster in this embodiment.
[0019] Figure 2 This is a flow chart of the non-intrusive resource allocation method based on Kubernetes cluster in this embodiment.
[0020] Figure 3 Schematic diagram of node resource allocation of the non-intrusive resource allocation method based on Kubernetes cluster in this embodiment.
[0021] Figure 4 This is a schematic diagram of the proxy server destruction process based on non-intrusive resource allocation of the Kubernetes cluster in this embodiment.
[0022] Figure 5 This is a flowchart of a GPUA resource allocation process based on non-intrusive resource allocation of a Kubernetes cluster in this embodiment.
[0023] Figure 6 This is a flowchart of the operation of the proxy server based on the non-intrusive resource allocation of the Kubernetes cluster in this embodiment.
[0024] Figure 7 This is a flowchart of another preferred non-intrusive resource allocation method based on a Kubernetes cluster in this embodiment.
[0025] Figure 8 This is a structural block diagram of a non-intrusive resource allocation device based on a Kubernetes cluster in this embodiment. DETAILED DESCRIPTION
[0026] In order to more clearly understand the purpose, technical solutions and advantages of the present application, the present application is described and illustrated below in conjunction with the accompanying drawings and embodiments.
[0027] Unless otherwise defined, the technical terms or scientific terms involved in this application shall have the general meaning understood by people with general skills in the technical field to which this application belongs. The words "one", "a", "the", "these" and the like in this application do not indicate a quantitative limitation, and they may be singular or plural. The terms "include", "comprise", "have" and any variants thereof involved in this application are intended to cover non-exclusive inclusions; for example, a process, method and system, product or device comprising a series of steps or modules (units) is not limited to the listed steps or modules (units), but may include unlisted steps or modules (units), or may include other steps or modules (units) inherent to these processes, methods, products or devices. The words "connect", "connected", "coupled" and the like involved in this application are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The "multiple" involved in this application refers to two or more. "And / or" describes the association relationship of associated objects, indicating that there may be three relationships, for example, "A and / or B" may mean: A exists alone, A and B exist at the same time, and B exists alone. Generally, the character " / " indicates that the objects associated with each other are in an "or" relationship. The terms "first", "second", "third", etc. in this application are only used to distinguish similar objects and do not represent a specific ordering of the objects.
[0028] The method embodiment provided in this embodiment can be executed in a terminal, a computer or a similar computing device. For example, running on a terminal, Figure 1 : is a hardware structure block diagram of a terminal of the non-intrusive resource allocation method based on a Kubernetes cluster in this embodiment. Figure 1 As shown, the terminal may include one or more ( Figure 1 Only one is shown in the figure) processor 102 and memory 104 for storing data, wherein processor 102 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA. The above terminal may also include a transmission device 106 and an input and output device 108 for communication functions. It can be understood by those skilled in the art that Figure 1 The structure shown is for illustration only and does not limit the structure of the above terminal. Figure 1 More or fewer components as shown, or with Figure 1 Different configurations shown.
[0029] The memory 104 can be used to store computer programs, for example, software programs and modules of application software, such as the computer program corresponding to the non-intrusive resource allocation method based on the Kubernetes cluster in the present embodiment. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, to implement the above method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include a memory remotely arranged relative to the processor 102, and these remote memories may be connected to the terminal via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0030] The transmission device 106 is used to receive or send data via a network. The above network includes a wireless network provided by a communication provider of the terminal. In one example, the transmission device 106 includes a network adapter (Network Interface Controller, referred to as NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 106 can be a radio frequency (Radio Frequency, referred to as RF) module, which is used to communicate with the Internet wirelessly.
[0031] In this embodiment, a resource allocation method is provided. Figure 2 is a flow chart of the non-intrusive resource allocation method based on Kubernetes cluster in this embodiment. Figure 2 As shown, the process includes the following steps: Step S201, in the K8s cluster node, after receiving the registration request of the GPU resource device plug-in, a proxy server for the GPU resource device plug-in is created according to preset configuration information.
[0032] Specifically, Kubernetes (often referred to as k8s) is an open source container orchestration platform for automating the deployment, expansion, and management of containerized applications. The k8s cluster consists of multiple nodes, divided into master nodes (Master Node) and worker nodes (WorkerNode), where nodes can be physical servers or virtual machines, which are machines running containerized applications. The master node is responsible for cluster management and contains components such as API Server, Scheduler, Controller Manager, and etcd. The worker node is responsible for running application containers. Each worker node runs the k8s main node agent (kubelet), which is responsible for managing the life cycle of the container. Pod is the smallest deployment unit in k8s and can contain one or more containers. Containers within a Pod share the same network namespace and storage volumes, so they can communicate with each other as if they were on the same machine. The controller is a component in k8s used to manage the life cycle and status of the Pod. Common controllers include: Deployment: used to manage stateless applications; StatefulSet: used to manage stateful applications; DaemonSet: Ensures that a Pod copy runs on each node; Job: used to run one-time tasks.
[0033] In a k8s cluster, GPU (Graphics Processing Unit) resources are a special hardware resource that is usually used to accelerate computationally intensive tasks such as machine learning, deep learning, scientific computing, and graphics rendering. Due to the particularity of GPU resources, k8s needs to manage and schedule GPU resources through special plug-ins or extensions. k8s itself does not directly manage GPU resources. The role of the GPU resource device plug-in is to expose GPU devices as manageable resources to the k8s cluster. The plug-in registers the GPU device to the k8s resource management system by interacting with the k8s Device Plugin API. Once registered, k8s can manage and schedule GPU resources like managing CPU and memory resources. The k8s scheduler decides which node to assign the Pod to run based on the Pod's resource request (including GPU resources). For example, if a Pod needs GPU resources, the scheduler will find a node with an available GPU and ensure that the GPU resources meet the Pod's needs. The GPU resource device plug-in will provide GPU status information (such as whether it is idle, model, etc.) to help the scheduler make decisions.
[0034] At present, when the scheduler finds an available GPU node, it searches based on the node. All GPU resources in the node are visible to k8s. K8s randomly selects GPUs from the available GPU nodes to allocate Pods. When there are many tasks, there is often a situation of resource competition. Therefore, before allocating Pods, this embodiment first deploys an agent service (Agent) and an agent service management program (ProxyManager) in each node of the k8s cluster, obtains GPU resources in the node through the Agent on each node and reports them to the ProxyManager through the GPU resource device plug-in. When the GPU resource device plug-in initiates a registration request to the ProxyManager program, the ProxyManager program creates a proxy server (Proxy) according to preset configuration information (such as the grouping strategy and hidden switch configuration of GPU resources, etc.). For example, a k8s cluster containing two types of GPU cards, namely GPUA and GPUB, is deployed on each node. Two types of GPU resource device plug-ins are installed in the cluster, namely DeviceA and DeviceB, which are used to report different types of GPU resources. The resource names registered in k8s are ResourceA and ResourceB. Figure 3 : is a schematic diagram of node resource allocation of the non-intrusive resource allocation method for the Kubernetes cluster of this embodiment, such as Figure 3As shown, there are 4 nodes, of which GPUB is arranged in two nodes, which are respectively recorded as ResourceB:5 and ResourceB:2; ResourceB:5 means that there are 5 GPUBs in the node, and ResourceB:2 means that there are 2 GPUBs in the node; GPUA is arranged in the other two nodes, which are respectively recorded as ResourceA:8 and ResourceA:8; similarly, ResourceA:8 means that there are 8 GPUA in the node. Taking the two nodes where GPUA is located as an example, the left figure shows the GPUA card in node 1, and the right figure shows the GPUA card in node 2. There are 8 GPU cards on each node, and the configuration information requires that the 4 cards on node 1 are configured to group Group1, and the other 4 cards on node 1 are configured to group Group2; there are 2 cards on node 2 configured to group Group1, and the switch of one card is set to off, and the other 6 cards on node 2 are configured to group Group2. According to the configuration information, before allocation, there are 8 GPUA cards in node 1, and the resource name registered in k8s is ResourceA:8; there are 8 GPUA cards in node 2, and the resource name registered in k8s is ResourceA:8; two proxy servers Proxy1 and Proxy2 are created on node 1, and two proxy servers Proxy3 and Proxy4 are also created on node 2. The created proxy servers are used to manage the GPU resources on each node respectively. After the proxy service, the reported resource specification names become Resource_Group1A and Resource_Group2A. The number of visible resources on the two nodes has also changed. Since a card is hidden on node 2, the number of GPU cards visible on node 2 has become 7.
[0035] During this process, a monitoring channel is established to continuously monitor changes in GPU resources on the node. Once changes in GPU resources are found, the GPU information cache module built into the Proxy is updated, and the resource information change notification is reported to the main node agent (Kubelet). The Proxy saves the service address of the corresponding GPU resource device plug-in, and forwards Kubelet's resource request to the GPU resource device plug-in; the Proxy receives the response result of the GPU resource device plug-in, customizes the response data according to the configuration of the extension point, and returns it to Kubelet.
[0036] In particular, when a user modifies the configuration of a GPU resource, ProxyManager uses Agent to determine whether there is a task in progress on the current GPU resource; if there is a task in progress, the configuration change fails; if there is no task in progress, the configuration information of the GPU resource is changed, and the Proxy to which the historical group related to the GPU resource belongs and the Proxy to which the change belongs are notified; the affected Proxy re-reports the GPU resource information to the Kubelet main node agent (Kubelet). This implements dynamic update of GPU resource configuration information.
[0037] Step S202: segmenting the GPU resources on the node through a proxy server, wherein the GPU resources on the node are obtained through a deployed proxy service.
[0038] Specifically, the GPU resources in the nodes are managed separately by deploying proxy servers in the nodes. For example, in the above step S201, two proxies are respectively set on node 1 and node 2. Figure 3 As shown, Proxy1 is used to manage four GPU cards in Node 1 ( Figure 3 The four GPUs in black text on a white background in the lower left image are named GPUA), and the resource name of the GPU card managed by Proxy1 is named Resource_Group1A and reported to the k8s cluster. The remaining four GPU cards in Node 1 are managed by Proxy2 ( Figure 3 The four GPUs in black and white on the lower left of the figure are named Resource_Group2A, and the resource name of the GPU card managed by Proxy2 is reported to the k8s cluster to establish a mapping relationship between the proxy server Proxy and the GPU resources. Similarly, in Node 2, Proxy3 manages the two GPU cards in Node 2 ( Figure 3 The two GPUs in the lower right corner of the picture are in black text on a white background. And one of the GPU cards is turned off ( Figure 3 The GPU card managed by Proxy3 is named Resource_Group1A and reported to the k8s cluster; the remaining 6 GPU cards in Node 2 are managed by Proxy4 ( Figure 3In the lower right image, the 6 GPUA with black background and white text are shown), and the resource name of the GPU card managed by Proxy4 is named Resource_Group2A and reported to the k8s cluster. In Node 1 and Node 2, the GPU resources in Node 1 and Node 2 are divided by creating 4 proxies. When the user needs to configure GPUA resources through the k8s cluster, according to the user's permissions, for example, the user can only view the GPU resources on the call group Group1, Node 1 is displayed to the k8s cluster as: Resource_Group1A: 4, Resource_Group2A: 4; Node 2 is displayed to the k8s cluster as: Resource_Group2A: 6, Resource_Group1A: 1, and the other GPUA is in a closed state and is not visible. At this time, the user can only view 4 GPU cards Resource_Group1A:4 in Group1 in Node 1 and 1 GPU card Resource_Group1A:1 in Group1 in Node 2 through the k8s cluster. For the user, Group2 in Node 1 and Node 2 is invisible at this time and can be viewed and called by users with other permissions. When the user's Pod requests GPUA resources, the resource identifier requested by the pod is switched from the original GPUA resource specification name to the Resource_Group1A resource specification name. The k8s scheduler will search for suitable GPUA cards in Group1 in Node 1 and Node 2 when scheduling the Pod, thereby achieving a finer division of resources in the node, achieving a more reasonable resource allocation, avoiding resource competition, and improving resource utilization.
[0039] Through the above steps S201 to S202, in the K8s cluster node, after receiving the registration request of the GPU resource device plug-in, a proxy server for the proxy GPU resource device plug-in is created according to the preset configuration information; the GPU resources on the node are divided by the proxy server, wherein the GPU resources on the node are obtained by the deployed proxy service. Compared with the resource management method in the prior art and the node-based resource management method, this embodiment deploys a proxy service and a proxy service management program in the node, and creates a proxy server in each node according to the preset configuration information through the proxy service management program to manage the GPU resources in each node respectively, thereby realizing the division of resources in the node, realizing further resource segmentation management, avoiding resource competition, realizing more rational resource allocation, and improving resource utilization.
[0040] In some of the embodiments, before receiving the registration request of the GPU resource device plug-in, the process further includes: initializing the configuration information of the GPU resources on the node.
[0041] Specifically, the GPU resource information in the node is obtained through the agent service Agent pre-deployed in the node to form an original configuration file. At this time, the configuration information of the GPU resources on all nodes in the original configuration file is in the default state, the grouping information and hidden switch configuration of the GPU resources on all nodes are initialized, and the original default state is cleared to facilitate subsequent user customized configuration.
[0042] In another embodiment, after the GPU resources on the node are divided through the proxy server, it also includes: executing corresponding configuration strategies according to preset configuration information through the interface extension point mechanism, wherein each configuration strategy corresponds to a configuration item function in the preset configuration information.
[0043] Specifically, after the GPU resources on the node are divided, this embodiment adopts an interface extension point mechanism to perform different configuration strategies according to the preset GPU resource configuration information. Among them, the extension point is a mechanism in the system design, which allows developers or users to customize the behavior of the system through the extension strategy, and add new functions or modify existing functions without changing the core logic of the system. Each configuration extension strategy usually corresponds to a configuration item function in the system configuration list, wherein the extension strategy can be a plug-in mechanism or a configuration driver, etc. By developing plug-ins or extension modules, the system can execute these plug-ins at runtime. The plug-in can process specific interface requests according to the functional requirements of the configuration item. Through the configuration items in the configuration file or database, the system can dynamically adjust the behavior according to these configurations. For example, the configuration item can specify which algorithm to use, enable or disable certain functions, etc., such as controlling the closing and opening of the GPU card. Different extension strategies can be set through the extension point mechanism, and the corresponding configuration information can be obtained according to the different extension strategies set. The operation and maintenance personnel can assemble their own operation and maintenance instructions by operating the configuration items, and the proxy service management program ProxyManager creates and manages the required proxy server Proxy according to the configuration information to proxy the GPU resources in each node. Through the extension point mechanism, the system can quickly adapt to new business needs through simple configuration or plug-in development without modifying the core code; by separating the core code from the extension logic, the system maintenance is easier, and the development and testing of extension modules are more independent; in addition, the system can dynamically load or unload extension modules as needed, supporting unlimited expansion. By reasonably designing extension points and expansion strategies, the scalability of the system and user experience can be significantly improved.
[0044] In some of the embodiments, the configuration item functions include grouping information of GPU resources and hidden switch configuration of GPU resources.
[0045] In the above embodiment, when the corresponding configuration strategy is executed through the extension point mechanism, each configuration strategy corresponds to a configuration item function in the preset configuration information, wherein the configuration item function mainly configures the grouping information of the GPU resources and the hidden switch of the GPU resources, that is, sets the number of groups of GPU resources in the node, sets the opening and closing of the GPU resources in the node, and controls the available GPU resources.
[0046] In another embodiment, a proxy server for proxy GPU resource device plug-ins is created according to preset configuration information, including: reading the preset configuration information, determining the number of groups of GPU resources according to the configuration information; and creating the same number of proxy servers for proxy GPU resource device plug-ins as the number of groups of GPU resources.
[0047] Specifically, the proxy service management program (ProxyManager) reads preset configuration information, wherein the preset configuration information includes the number of groups and switch status of GPU resources, creates the same number of proxy servers Proxy according to the number of groups of GPU resources in the configuration information, analyzes the GPU resource grouping on the node, creates multiple proxy servers Proxy if there are multiple groups, and establishes a mapping relationship between Proxy and GPU resources.
[0048] In some of the embodiments, after creating a number of proxy servers for proxy GPU resource device plug-ins that are the same as the number of groups of GPU resources, it also includes: registering the resource specification name with the main node agent on the node through the proxy server to obtain a new resource specification name, so that k8s can call GPU resources according to the new resource specification name.
[0049] After the proxy server Proxy is created according to the number of groups of GPU resources in the configuration information, each proxy server Proxy registers the new resource specification name after the proxy to the main node agent (Kubelet) on the node, such as Resource_Group1A and Resource_Group2A described in the above step S202, which are the resource specification names of the new proxy GPU resources registered by each proxy server Proxy to the k8s cluster. Among them, Kubelet is a core component in k8s, running on each node in the cluster. It is the main node agent of the node, managing the containers and Pods on the node, and ensuring that they run correctly according to the instructions of the k8s control plane (Control Plane). Kubelet is a bridge between the node and the control plane in the k8s cluster, responsible for performing node-related operations. By reporting the new resource specification names of each proxy server Proxy, k8s can call new GPU resources according to user needs.
[0050] In another embodiment, when the GPU resource device plug-in on the node is offline, the proxy server corresponding to the GPU resource device plug-in is destroyed.
[0051] Based on the management of the entire life cycle of the proxy server, the Proxy is created when the ProxyManager receives the registration of the GPU resource device plug-in; when the GPU resource device plug-in is detected to be offline, uninstalled or crashed, the connection between the Proxy and the GPU resource device plug-in is cut off, and the Proxy is automatically destroyed through the ProxyManager. This avoids affecting the normal operation of the k8s cluster and saves the storage resources of the k8s cluster. Figure 4 FIG. 4 is a schematic diagram of a proxy server destruction process based on non-intrusive resource allocation of a Kubernetes cluster in this embodiment. Figure 4 As shown, in the process of the GPU resource device plug-in service performing proxy service through the proxy server (Proxy), a monitoring channel is established to continuously monitor the GPU resource changes on the node. When it is detected that the GPU resource device plug-in channel is disconnected, the connection between the Proxy and the GPU resource device plug-in is cut off, and the proxy server (Proxy) is automatically destroyed through the proxy service management program (ProxyManager).
[0052] For example, Figure 5 is a flowchart of GPUA resource allocation based on non-intrusive resource allocation of Kubernetes cluster in this embodiment, such as Figure 5 As shown in the figure, first, in the node of the k8s cluster, the GPU resource device plug-in (DeviceA) initiates a registration request to the proxy service manager (ProxyManager); secondly, the proxy service manager (ProxyManager) loads the configuration list and creates two proxy servers (ProxyDeviceAGroup1 and ProxyDeviceAGroup2) according to the configuration list, and divides the GPU resources in the node into two groups through the proxy server (Proxy); finally, the proxy server (Proxy) registers the new resource specification names (Resource_Group1A and Resource_Group2A) of the divided GPU resources to the main node agent (Kubelet) on the node to complete the allocation of GPUA resources in the node.
[0053] For example, Figure 6 This is a flow chart of the operation of the proxy server based on the non-intrusive resource allocation of the Kubernetes cluster in this embodiment. Figure 6 As shown, the process includes: Step 1: Register the service. The GPU device plug-in first registers the service with the proxy service manager (ProxyManager), indicating that it can manage GPU device resources so as to provide remote procedure call framework services (grpc services). Step 2: Generate a Proxy component. After receiving the registration request, ProxyManager generates one or more Proxy components according to the configuration file for communication between the GPU device plug-in and kubelet. Step 3: Proxy registers the plug-in (device) resource with the primary node agent (Kubelet) and provides a Unix domain socket (sock) address for subsequent communication; Step 4: Initiate a request. When there is a request to access the GPU device, kubelet initiates a request to the Proxy proxy to obtain GPU device information or perform related operations. Step 5: Forward the request. The Proxy agent forwards the kubelet request to the GPU device plug-in service. Step 6: Receive the plug-in response information. The plug-in service receives and processes the request from the Proxy service, obtains the relevant information of the GPU device, and the Proxy service receives the plug-in response information. Step 7: Intercept the response body. During the request transmission process, the interceptor will intercept and process the plug-in response information; Step 8: Return the response result. The Proxy proxy service returns the response result of the plug-in to kubelet.
[0054] In the k8s node, kubelet can dynamically obtain and manage GPU device resources through the Proxy service. The management of GPU resources is further refined. The Proxy acts as a bridge in this process, responsible for forwarding requests and delivering responses, making the management of GPU resources more flexible and efficient, and also easier to expand and maintain.
[0055] This embodiment also provides a resource allocation method. Figure 7 is a flowchart of another preferred non-intrusive resource allocation method based on a Kubernetes cluster in this embodiment. Figure 7 As shown, the process includes the following steps: Step S701, deploying a proxy service and a proxy service management program in each node of the K8s cluster; Step S702, obtaining GPU resources on the node through the proxy service, and initializing the configuration information of the GPU resources on the node; Step S703, deploying a GPU resource device plug-in on the node, and the GPU resource device plug-in initiates a registration request to the proxy service management program; Step S704, the proxy service management program determines the number of groups of GPU resources according to preset configuration information, and creates proxy servers of proxy GPU resource device plug-ins of the same number as the number of groups of GPU resources, and divides the GPU resources on the node through the proxy servers; Step S705, registering the resource specification name with the main node agent on the node through the proxy server, and obtaining a new resource specification name, so that k8s can call GPU resources according to the new resource specification name; Step S706: Execute the corresponding configuration strategy according to the retrieved configuration information of the GPU resources through the interface extension point mechanism.
[0056] Through the above steps S701 to S706, compared with the resource management method in the prior art and the node-based method, this embodiment obtains all GPU resource information on the node through the proxy service deployed on the node, and the GPU resource initiates a registration request to the proxy service manager. The proxy service manager creates a GPU resource device plug-in corresponding to the proxy server agent according to the preset configuration information, thereby realizing the segmentation of GPU resources on the node, obtaining the segmented GPU resources, and then scheduling the GPU resources according to the native scheduling mechanism of k8s. Among them, by establishing different proxy servers on the node to proxy different GPU resources, the resource cutting and allocation is realized on the node, and the resources are subdivided on the node to improve the utilization of resources.
[0057] In this embodiment, a resource allocation device is also provided, which is used to implement the above-mentioned embodiments and preferred implementation modes, and the descriptions that have been made will not be repeated. The terms "module", "unit", "subunit" and the like used below can implement a combination of software and / or hardware for a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceivable.
[0058] Figure 8 is a structural block diagram of a non-intrusive resource allocation device based on a Kubernetes cluster in this embodiment, such as Figure 8 As shown, the device 80 includes: a creation module 81 and a segmentation module 82, wherein: A creation module 81 is used to create a proxy server for the GPU resource device plug-in according to preset configuration information in a K8s cluster node after receiving a registration request of the GPU resource device plug-in; The segmentation module 82 is used to segment the GPU resources on the node through the proxy server, wherein the GPU resources on the node are obtained through the deployed proxy service.
[0059] In this embodiment, an electronic device is further provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.
[0060] Optionally, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor, and the input / output device is connected to the processor.
[0061] Optionally, in this embodiment, the processor may be configured to perform the following steps through a computer program: S1, in a K8s cluster node, after receiving a registration request of a GPU resource device plug-in, a proxy server for the GPU resource device plug-in is created according to preset configuration information; S2, dividing the GPU resources on the node through the proxy server, wherein the GPU resources on the node are obtained through the deployed proxy service.
[0062] It should be noted that the specific examples in this embodiment can refer to the examples described in the above embodiments and optional implementation modes, and will not be repeated in this embodiment.
[0063] In addition, in combination with the resource allocation method provided in the above embodiment, a storage medium may be provided in this embodiment to implement the method. The storage medium stores a computer program; when the computer program is executed by a processor, any non-intrusive resource allocation method based on a Kubernetes cluster in the above embodiment is implemented.
[0064] It should be understood that the specific embodiments described herein are only used to explain the application, rather than to limit it. Based on the embodiments provided in this application, all other embodiments obtained by ordinary technicians in this field without creative work are within the protection scope of this application.
[0065] Obviously, the drawings are only some examples or embodiments of the present application. For ordinary technicians in the field, the present application can also be applied to other similar situations based on these drawings without creative work. In addition, it is understandable that although the work done in this development process may be complicated and lengthy, for ordinary technicians in the field, certain changes in design, manufacturing or production based on the technical content disclosed in this application are only conventional technical means and should not be regarded as insufficient content disclosed in this application.
[0066] The term "embodiment" in this application refers to a specific feature, structure or characteristic described in conjunction with the embodiment that can be included in at least one embodiment of the present application. The appearance of this phrase in various locations in the specification does not necessarily mean the same embodiment, nor does it mean that it is mutually exclusive with other embodiments and is independent or optional. It is clearly or implicitly understood by those of ordinary skill in the art that the embodiments described in this application can be combined with other embodiments without conflict.
[0067] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. Non-relational databases may include distributed databases based on blockchains, etc., but are not limited to this. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., but are not limited to this.
[0068] The above-mentioned embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of patent protection. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the scope of protection of the present application. Therefore, the scope of protection of the present application shall be subject to the attached claims.
Claims
1. A non-intrusive resource allocation method based on Kubernetes cluster, characterized in that: include: In the K8s cluster node, after receiving the registration request of the GPU resource device plug-in, a proxy server for the GPU resource device plug-in is created according to the preset configuration information; The GPU resources on the node are divided by the proxy server, wherein the GPU resources on the node are obtained by the deployed proxy service.
2. The non-intrusive resource allocation method based on Kubernetes cluster according to claim 1 is characterized in that: Before receiving the registration request of the GPU resource device plug-in, the method further includes: Initialize the configuration information of the GPU resources on the node.
3. The non-intrusive resource allocation method based on Kubernetes cluster according to claim 1 is characterized in that: After the GPU resources on the node are divided by the proxy server, the method further includes: The corresponding configuration strategy is executed according to the preset configuration information through the interface extension point mechanism, wherein each configuration strategy corresponds to a configuration item function in the preset configuration information.
4. The non-intrusive resource allocation method based on Kubernetes cluster according to claim 3 is characterized in that: The configuration item functions include grouping information of the GPU resources and hidden switch configuration of the GPU resources.
5. The non-intrusive resource allocation method based on Kubernetes cluster according to claim 1, characterized in that: The step of creating a proxy server for the GPU resource device plug-in according to the preset configuration information includes: Reading the preset configuration information, and determining the number of groups of the GPU resources according to the configuration information; Create a number of proxy servers that are proxies for the GPU resource device plug-ins that is the same as the number of groups of the GPU resources.
6. The non-intrusive resource allocation method based on Kubernetes cluster according to claim 5 is characterized in that: After creating the same number of proxy servers as the number of groups of the GPU resources that represent the GPU resource device plug-ins, the method further includes: The proxy server registers the resource specification name with the main node agent on the node to obtain a new resource specification name, so that the k8s cluster can retrieve GPU resources according to the new resource specification name.
7. The non-intrusive resource allocation method based on Kubernetes cluster according to claim 1, characterized in that: The method further comprises: When the GPU resource device plug-in on the node is offline, the proxy server corresponding to the GPU resource device plug-in is destroyed.
8. A non-intrusive resource allocation device based on Kubernetes cluster, characterized in that: include: Create modules and split modules, where The creation module is used to create a proxy server that acts as a proxy for the GPU resource device plug-in in the K8s cluster node according to preset configuration information after receiving a registration request for the GPU resource device plug-in; The segmentation module is used to segment the GPU resources on the node through the proxy server, wherein the GPU resources on the node are obtained through the deployed proxy service.
9. An electronic device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to run the computer program to execute the non-intrusive resource allocation method based on a Kubernetes cluster according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the non-intrusive resource allocation method based on a Kubernetes cluster described in any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
GPU resource-based data processing method and system, and electronic equipment
CN110764901A
GPU allocation method and system, storage medium and equipment
CN112463383A
Resource scheduling method and device, computer equipment and storage medium
CN115145695A
GPU resource scheduling method and device, electronic equipment and storage medium
CN115373861A
GPU computing resource management method and device, electronic equipment and readable storage medium
CN115562878A
Cited By
Implementation method and device for creating load by specified GPU, equipment and storage medium
CN121233349A
Method, device and storage medium for creating a load implementation method by designating a GPU
CN121233349B