Heterogeneous resource optimization method and device under Pod capacity reduction scene based on K8s
By calculating the GPU topology distribution of Pod nodes in Kubernetes system and optimizing the Pod scaling process, the problem of GPU resource fragmentation is solved and the utilization rate of GPU resources is improved.
Patent Information
- Application Number
- CN202510285311.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-06-10
AI Technical Summary
The existing Kubernetes system fails to fully consider the characteristics of GPU resources when reducing the pod, resulting in the fragmentation and low utilization of GPU resources, especially in the mixed use scenarios of multi-node, multi-service, and multi-spec GPU resources.
A heterogeneous resource optimization method in the K8s-based Pod scaling scenario is proposed. By calculating the GPU topology distribution of all Pod nodes under the target service, the spare GPU resources are determined, and more GPU cards are vacancies when scaling is reduced, so as to reduce resource fragmentation.
Effectively reduce the degree of fragmentation of GPU resources and improve the utilization rate of GPU resources, especially in a mixed scenario of multiple resource specifications and services.
Smart Images

Figure CN120123045A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of cloud platforms, and specifically provides a method and device for optimizing heterogeneous resources in a Pod scaling scenario based on K8s. Background Art
[0002] Kubernetes (K8s) is the most popular container orchestration system today, which can automate the deployment, expansion and management of Pod (the smallest unit of work in K8s). In K8s, the matching of Pod nodes and Node nodes (physical machines or virtual machines) is the responsibility of Scheduler, while the management of the number of copies of Pod nodes is the responsibility of ReplicaSet.
[0003] However, with the widespread application of hardware accelerators such as GPUs in deep learning, graphics processing and other fields, more and more applications need to run GPU workloads on Kubernetes. However, the existing ReplicaSet controller does not fully consider the particularity of GPU, a scarce hardware resource. For example, when scaling down Pods, the default scaling strategy does not take into account the characteristics and requirements of GPU resources, which often leads to fragmentation and low utilization of GPU resources.
[0004] Currently, the sorting algorithm used by K8s for Pod scaling operations is as follows:
[0005] Unstarted Pod nodes take priority.
[0006] Pod nodes with more Pods in the same replicaset will be prioritized over those with fewer Pods.
[0007] Pods created later take precedence.
[0008] K8s's own Pod recycling strategy tends to be stable. Except for the highest priority of recycling Pods that cannot provide services as much as possible, the other two strategies tend to make the number of Pods on each node as even as possible, and make the Pods that have already provided services continue to provide services as much as possible. This is reasonable for normal stability, but for resources such as GPUs, which have extremely limited quantity, multiple relatively independent devices on a single node, and are not highly compressible, the average strategy will inevitably cause resource fragmentation during recycling. Especially in the mixed use scenario of multi-node, multi-service, and multi-specification GPU resources, the default shrinking strategy may lead to fragmentation of GPU resources.
[0009] In view of this, the present invention patent is proposed. Summary of the invention
[0010] In view of the above technical problems, the present invention proposes a heterogeneous resource optimization method and device based on K8s in a Pod scaling scenario, which is used to recycle GPU resources when scaling down or iterating GPU Deploy, so as to free up more GPU resources as much as possible.
[0011] Specifically, the following technical solutions are adopted:
[0012] In a first aspect, the present invention provides a heterogeneous resource optimization method in a Pod scaling scenario based on K8s, comprising:
[0013] When the target service is scaled down, the GPU topology distribution of all Pod nodes under the target service is calculated;
[0014] Based on the GPU topology distribution, the available GPU resources are determined and the Pod is scaled down.
[0015] As an optional implementation of the present invention, a heterogeneous resource optimization method in a Pod scaling scenario based on K8s of the present invention includes:
[0016] By using the webhook capability in K8s, the Rs-webhook component is provided;
[0017] When the replicaset resources of K8s are changed, the corresponding interface of the Rs-webhook component is called, and the replicaset resources before and after the change are passed in as parameters;
[0018] The Rs-webhook component determines the replicas field value before and after the replicaset resource of K8s is changed, and determines whether the target service is scaled down.
[0019] As an optional implementation of the present invention, in a heterogeneous resource optimization method based on a K8s Pod scaling scenario of the present invention, the Rs-webhook component determines the replicas field value before and after the K8s replicaset resource is changed, and determines whether the target service is scaled down, including:
[0020] If the replicas field value of the replicaset resource before the change is greater than the replicas field value of the replicaset resource after the change, it is determined that the target service is being scaled down; otherwise, it is determined that the target service is not being scaled down.
[0021] As an optional embodiment of the present invention, in a heterogeneous resource optimization method in a Pod scaling scenario based on K8s of the present invention, the GPU topology distribution of all Pod nodes under the computing target service includes:
[0022] The Rs-webhook component obtains the number N of Pod nodes that have been scaled down, obtains all Pod nodes under the replicas field value of the replicaset resource before the change, and the Node nodes corresponding to the Pod nodes;
[0023] For any Node node, obtain the number of idle GPUs G1 corresponding to the Node node, obtain the number of Pod nodes N1 on the Node node that belongs to the replicaset resource that is currently being scaled down, and calculate the number of available GPUs G2 and the number of Node nodes that can be scaled down N2;
[0024] Calculate G2=G1+N2 and traverse all Node nodes;
[0025] Get the available GPU quantity G2 and the scalable capacity N2 of all Node nodes, and sort all Pod nodes.
[0026] As an optional implementation of the present invention, in a heterogeneous resource optimization method based on K8s in a Pod scaling scenario of the present invention, for any Node node, obtaining the number of idle GPUs G1 corresponding to the Node node, obtaining the number of Pod nodes N1 on the Node node belonging to the replicaset resource currently being scaled down, and calculating the number of available GPUs G2 and the number of Node nodes that can be scaled down N2 include:
[0027] When the scaled-down replicaset resource is an exclusive service, the calculation condition is: if N>N1, then let N2=N1; otherwise, let N2=N.
[0028] As an optional implementation of the present invention, in a heterogeneous resource optimization method based on K8s in a Pod scaling scenario of the present invention, for any Node node, obtaining the number of idle GPUs G1 corresponding to the Node node, obtaining the number of Pod nodes N1 on the Node node belonging to the replicaset resource currently being scaled down, and calculating the number of available GPUs G2 and the number of Node nodes that can be scaled down N2 include:
[0029] When the replicaset resource being scaled down is a shared service, the calculation condition is as follows: traverse all GPUs of any Node node, obtain the number of Pod nodes N3 on the GPU that belong to the replicaset resource currently being scaled down, calculate whether the GPU has free space G3 and the number of Pod nodes N4 that can be scaled down on the GPU, and the calculation condition is as follows: when N>N3, set N4=N3, otherwise N4=N;
[0030] Calculate G3 = the number of Pod nodes on the GPU – N4, and store the relationship among GPU, pod, and G3 in memory.
[0031] As an optional embodiment of the present invention, in a heterogeneous resource optimization method based on K8s in a Pod scaling scenario of the present invention, obtaining the number of available GPUs G2 and the number of scaling N2 of all Node nodes, and sorting all Pod nodes includes:
[0032] When the replicaset resource to be scaled down is an exclusive service, obtain the available GPU quantity G2 and the scalable quantity N2 of all Node nodes, and sort them in descending order of G2 and ascending order of N2;
[0033] When the scaled-down replicaset resource is a shared service, the value corresponding to G3 is obtained from the memory, and the Pod nodes with G3 of 0 are filtered out. The Pod node list is divided into two groups, 0 and non-0, and arranged in reverse order of G2 and in ascending order of N2 respectively. The result list is merged into one list in the order of group 0 and group non-0.
[0034] As an optional embodiment of the present invention, in a heterogeneous resource optimization method in a Pod scaling scenario based on K8s of the present invention, the determining of the free GPU resource situation based on the GPU topology distribution and scaling down the Pod includes:
[0035] According to the sorting results of all Pod nodes, N Pod nodes are selected from the beginning, and the kube-apiserver component of K8s is called to mark the pod-deleted-cost as -1. The process ends and returns success.
[0036] In a second aspect, the present invention provides a heterogeneous resource optimization device in a Pod scaling scenario based on K8s, comprising:
[0037] GPU resource calculation module, when the target service is scaled down, calculates the GPU topological distribution of all Pod nodes under the target service;
[0038] The Pod reduction module determines the free GPU resources based on the GPU topology distribution and reduces the Pod capacity.
[0039] In a third aspect, the present invention provides a computer-readable recording medium storing a computer executable program. When the computer executable program is executed, the heterogeneous resource optimization method in the K8s-based Pod scaling scenario is implemented.
[0040] Compared with the prior art, the present invention has the following beneficial effects:
[0041] The present invention provides a heterogeneous resource optimization method based on K8s in a Pod scaling scenario, and provides a recycling solution dedicated to GPU resource services, which reduces resource fragmentation by giving priority to freeing up more GPU cards. Specifically, when the service is scaled down, by calculating the GPU topological distribution of all PODs under the entire service, the location with the most free GPU resources and only a smaller number of shrinkages is preferentially found for scaling down. In this way, when the pod is scaled down, the present invention can give priority to freeing up more GPU cards through a heterogeneous resource optimization method based on K8s in a Pod scaling scenario, effectively reducing the degree of GPU resource fragmentation. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 An example flow chart of a method for optimizing heterogeneous resources in a Pod scaling scenario based on K8s according to the first embodiment of the present invention;
[0043] Figure 2 An example diagram of Pod sorting of a heterogeneous resource optimization method in a Pod scaling scenario based on K8s according to the first embodiment of the present invention;
[0044] Figure 3 A schematic structural diagram of an electronic device according to a second embodiment of the present invention;
[0045] Figure 4 A schematic diagram of a computer-readable recording medium according to a second embodiment of the present invention. DETAILED DESCRIPTION
[0046] To make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be described clearly and completely in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them.
[0047] Therefore, the following detailed description of the embodiments of the present invention is not intended to limit the scope of the invention claimed for protection, but merely represents some embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0048] It should be noted that, in the absence of conflict, the embodiments of the present invention and the features and technical solutions in the embodiments may be combined with each other.
[0049] It should be noted that similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, further definition and explanation thereof is not required in subsequent drawings.
[0050] In the description of the present invention, it should be noted that the orientation or positional relationship indicated by terms such as "upper" and "lower" is based on the orientation or positional relationship shown in the drawings, or the orientation or positional relationship in which the product of the present invention is usually placed during use, or the orientation or positional relationship commonly understood by those skilled in the art. Such terms are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation to the present invention. In addition, terms such as "first" and "second" are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.
[0051] Embodiment 1
[0052] An optimization method for heterogeneous resources in the scenario of Pod scaling down based on K8s in this embodiment includes:
[0053] When the target service is scaled down, calculate the GPU topology distribution of all Pod nodes under the target service;
[0054] Based on the GPU topology distribution, determine the available GPU resources and perform Pod scaling down.
[0055] An optimization method for heterogeneous resources in the scenario of Pod scaling down based on K8s in this embodiment provides a recovery solution dedicated to GPU resource services, and reduces the degree of resource fragmentation by preferentially freeing up more GPU cards. Specifically, when the service is scaled down, by calculating the GPU topology distribution of all PODs under the entire service, preferentially find the position that can free up the most GPU resources and only requires scaling down a smaller number of Pods for scaling down. In this way, when the Pod is scaled down, through the optimization method for heterogeneous resources in the scenario of Pod scaling down based on K8s in this embodiment, more GPU cards can be preferentially emptied, effectively reducing the degree of GPU resource fragmentation.
[0056] An optimization method for heterogeneous resources in the scenario of Pod scaling down based on K8s in this embodiment includes:
[0057] By using the ability of webhook in K8s, provide the Rs-webhook component;
[0058] When there is a change in the replicaset resource of K8s, call the corresponding interface of the Rs-webhook component and pass the replicaset resources before and after the change as parameters;
[0059] The Rs-webhook component judges the value of the replicas field in the replicaset resource of K8s before and after the change to judge whether the target service is scaled down.
[0060] This embodiment uses the capabilities of webhooks in K8s to provide an Rs-webhook component and sets the ValidatingWebhookConfiguration to bind the UPDATE operation of the replicaset resource to the Rs-webhook service interface.
[0061] ValidatingWebhookConfig is an API resource in Kubernetes used to configure validation webhooks. A webhook is an HTTP callback mechanism that allows custom logic to be executed when specific events (such as creating, updating, or deleting resource objects) occur on the Kubernetes API server. Validating Webhooks are mainly used to validate resource objects before they are persisted to etcd to ensure they comply with specific rules or policies.
[0062] The role and importance of ValidatingWebhookConfig:
[0063] · Request validation: By configuring validation webhooks, requests submitted to the Kubernetes cluster can be validated to ensure that aspects such as the structure, format, content, and security of the requests meet expectations.
[0064] · Implementing policy rules: Administrators can create and configure validation rules to validate requests as needed. These rules can be defined based on the type, annotations, labels, or other attributes of the resource objects.
[0065] · Enhancing security: By using validation webhooks, malicious or invalid requests can be prevented from being submitted to the Kubernetes cluster, thereby improving the security of the cluster.
[0066] · Centralized management: Provides a mechanism for centralized management and configuration of validation webhooks. Administrators can centrally manage the request validation logic across the entire cluster by creating and managing ValidatingWebhookConfiguration objects.
[0067] How to configure and use ValidatingWebhookConfig in Kubernetes:
[0068] 1. Write the Webhook service
[0069] First, an HTTP service needs to be written to handle Webhook requests. This service can be a simple RESTful API that receives Webhook requests sent by Kubernetes and performs corresponding validation operations based on the event information in the requests.
[0070] 2. Deploy the Webhook service
[0071] Deploy the written Webhook service to the Kubernetes cluster to ensure that it can receive and process Webhook requests normally. Deployment, Pod, or other appropriate Kubernetes resource objects can be used to deploy the service.
[0072] 3. Create a ValidatingWebhookConfiguration resource object
[0073] Create a ValidatingWebhookConfiguration resource object to define the Webhook configuration. In this configuration, information such as the address of the Webhook service, the formats of requests and responses, and the types of events that trigger the Webhook needs to be specified.
[0074] Based on the K8s cluster, when there are changes to the replicaset resource (when the deploy is scaled down or updated, the K8s deploy controller will operate on the replicas resource. Scaling up or down means modifying the replicas field value of the current replicaset. When there is a change, a new replicaset will be generated to replace the old replicaset resource, and the replicas field value of the old replicaset will be scaled down, while the replicas field value of the new replicaset will be scaled up), the kube-apiserver will call the corresponding interface of the Rs-webhook component and pass the replicaset resources before and after the change as parameters.
[0075] The Rs-webhook component will determine whether the replicaset uses GPU resources and the replicas values before and after the change. If it does not use GPU resources or is scaling up, no action will be taken. If it is scaling down, it will enter the processing flow of the heterogeneous resource optimization method in the Pod scaling-down scenario.
[0076] Therefore, the Rs-webhook component determines the value of the replicas field before and after the change of the replicaset resource in K8s, determines whether the target service is scaled in, and further determines whether to enter the processing flow of a heterogeneous resource optimization method in the Pod scaling-in scenario based on K8s in this embodiment.
[0077] As an alternative implementation of this embodiment, in a heterogeneous resource optimization method in the Pod scaling-in scenario based on K8s in this embodiment, the Rs-webhook component determines the value of the replicas field before and after the change of the replicaset resource in K8s, and determining whether the target service is scaled in includes:
[0078] If the value of the replicas field of the replicaset resource before the change is greater than the value of the replicas field of the replicaset resource after the change, it is determined that the target service is scaled in; otherwise, it is determined that the target service is not scaled in.
[0079] In a heterogeneous resource optimization method in the Pod scaling-in scenario based on K8s in this embodiment, the Rs-webhook component determines the value of the replicas field before and after the change of the replicaset resource in K8s. If it is scaled in, it enters the processing flow of the heterogeneous resource optimization method in the Pod scaling-in scenario; if it is not scaled out, no processing is performed.
[0080] In addition, the Rs-webhook component described in this embodiment also determines whether the replicaset uses GPU resources. If it does not use GPU resources, no processing is performed. If it is scaled in, it enters the processing flow of the heterogeneous resource optimization method in the Pod scaling-in scenario.
[0081] Specifically, in a heterogeneous resource optimization method in the Pod scaling-in scenario based on K8s in this embodiment, calculating the GPU topology distribution of all Pod nodes under the target service includes:
[0082] The Rs-webhook component obtains the number N of Pod nodes to be scaled in, obtains all Pod nodes corresponding to the value of the replicas field of the replicaset resource before the corresponding change, and the Node nodes corresponding to the Pod nodes;
[0083] For any Node node, obtain the number G1 of idle GPUs on this Node node, obtain the number N1 of Pod nodes on this Node node that belong to the replicaset resource currently being scaled in, and calculate the available free GPU number G2 and the number N2 of nodes that can be scaled in on this Node node;
[0084] Calculate G2 = G1 + N2, and traverse all Node nodes;
[0085] Obtain the available free GPU quantity G2 and the reducible quantity N2 of all Node nodes, and sort all Pod nodes.
[0086] A heterogeneous resource optimization method based on K8s in the Pod scaling-down scenario of this embodiment, by obtaining the available free GPU quantity G2 and the reducible quantity N2 of all Node nodes, sorting all Pod nodes, when scaling down Pods, more GPU cards can be emptied preferentially, effectively reducing the fragmentation degree of GPU resources.
[0087] As an alternative implementation manner of this embodiment, in a heterogeneous resource optimization method based on K8s in the Pod scaling-down scenario of this embodiment, for any Node node, obtain the corresponding free GPU quantity G1 on this Node node, obtain the number N1 of Pod nodes belonging to the replicaset resource that is currently being scaled down on this Node node, and calculating the available free GPU quantity G2 and the reducible quantity N2 of this Node node includes:
[0088] When the replicaset resource being scaled down is an exclusive service, the calculation condition is: judge if N > N1, then let N2 = N1; otherwise let N2 = N.
[0089] As an alternative implementation manner of this embodiment, in a heterogeneous resource optimization method based on K8s in the Pod scaling-down scenario of this embodiment, for any Node node, obtain the corresponding free GPU quantity G1 on this Node node, obtain the number N1 of Pod nodes belonging to the replicaset resource that is currently being scaled down on this Node node, and calculating the available free GPU quantity G2 and the reducible quantity N2 of this Node node includes:
[0090] When the replicaset resource being scaled down is a shared service, the calculation condition is: traverse all GPUs of any Node node, obtain the number N3 of Pod nodes belonging to the replicaset resource that is currently being scaled down on this GPU, calculate whether this GPU is available for free G3 and the reducible quantity N4 on this GPU, and the calculation condition is as follows: when N > N3, let N4 = N3, otherwise N4 = N;
[0091] Calculate G3 = the number of Pod nodes on this GPU - N4, and store the relationship of GPU, pod, G3 in the memory.
[0092] Furthermore, in a heterogeneous resource optimization method based on K8s in the Pod scaling-down scenario of this embodiment, the obtaining the available free GPU quantity G2 and the reducible quantity N2 of all Node nodes and sorting all Pod nodes includes:
[0093] When the scaled-down replicaset resource is an exclusive service, obtain the available free GPU quantities G2 and the scalable quantity N2 of all Node nodes, and sort them in descending order of G2 and ascending order of N2.
[0094] When the scaled-down replicaset resource is a shared service, obtain the corresponding value of G3 from the memory, filter out the Pod nodes with G3 equal to 0, divide the Pod node list into two groups: 0 and non-0, sort them respectively in descending order of G2 and ascending order of N2, and merge the result lists into one list in the order of the 0 group and the non-0 group.
[0095] Further, in a heterogeneous resource optimization method for Pod scaling-down scenarios based on K8s in this embodiment, determining the available GPU resources based on the GPU topology distribution and performing Pod scaling-down includes:
[0096] According to the sorting results of all Pod nodes, select N Pod nodes from the beginning, call the kube-apiserver component of K8s to mark pod-deleted-cost as -1, and the process ends and returns successfully.
[0097] See Figure 1 and Figure 2 As shown in
[0098] Step S1, the Rs-webhook component compares and determines the replicaset before and after the change: When replicas before the change > replicas after the change, it is determined as scaling down, and continue to execute the following steps, otherwise terminate the process and directly return successfully.
[0099] Step S2, the Rs-webhook component obtains the number of Pods N to be scaled down, obtains all component Pods under the corresponding replicas, and the Node corresponding to the Pod. And traverse all Nodes and execute steps S3 - S6.
[0100] Step S3, obtain the corresponding free GPU quantity G1 on this Node, obtain the number of Pods N1 on this Node that belong to the replicaset currently being scaled down, calculate the available free GPU quantity G2 and the scalable quantity N2 of this node, and the calculation methods are as follows:
[0101] Step S3.1, when the scaled-down replicaset is an exclusive service, the calculation conditions are as follows: Determine that if N > N1, then let N2 = N1; otherwise let N2 = N.
[0102] Step S3.2, when the scaled-down replicaset is a shared service, calculate the conditions: Traverse all GPUs on this Node, obtain the number of Pods N3 affiliated with the current replicaset on this GPU, calculate whether the current GPU can be free G3 and the number N4 that can be scaled down on this card. The calculation conditions are as follows: When N > N3, let N4 = N3; otherwise, N4 = N. Calculate G3 = the number of Pods on this GPU - N4, and store the relationship between the GPU, Pod, and G3 in memory. The G2 of this node is equal to the number of GPUs with all G3 = 0 selected, and the N2 of this node is the number of Pods affiliated with this replicaset corresponding to all GPUs with G3 = 0 selected.
[0103] Step S4, finally calculate G2 = G1 + N2. After traversing all nodes, enter the next step.
[0104] Step S5, obtain the available free GPU number G2 and the number N2 that can be scaled down of all Nodes, and sort all Pods according to the following rules:
[0105] Step S5.1, if it is an exclusive service Pod, obtain the G2 and N2 of the node where it is located, and sort it in descending order of G2 and ascending order of N2;
[0106] Step S5.2, if it is a shared service, obtain the corresponding value of G3 from memory, filter out the Pods with G3 = 0, divide the Pod list into two groups of 0 and non-0, sort them in descending order of G2 and ascending order of N2 respectively, and merge the result lists into one list in the order of the 0 group and the non-0 group.
[0107] Step S6, select N Pods from the beginning of the result and call kube-apiserver to mark pod-deleted-cost as -1. The process ends and returns success.
[0108] When the service is scaled down, through a heterogeneous resource optimization method based on K8s in this embodiment, calculate the GPU topology distribution of all Pods under the entire service, and preferentially find the position that can free the most GPU resources and only needs to scale down a smaller number for scaling down. As Figure 2 shown, the scaling-down plan that can free 4 cards is better than the scaling-down plan that can free 3 + 1, because in the case of the same total amount, the former has one more allocation method that can accommodate 4-card Pods than the latter, and the number of fragments is less.
[0109] Therefore, an optimization method for heterogeneous resources in the Pod scale-down scenario based on K8s in this embodiment calculates the priority order of Pods in real time through the Rs-webhook component when the replicaset is scaled down and controls it through the pod-deleted-cost; by vacating more resources and reducing fragmentation, the allocation rate of GPU resources, especially in the scenario of mixed services with multiple resource specifications, is improved.
[0110] This embodiment also provides an optimization device for heterogeneous resources in the Pod scale-down scenario based on K8s, including:
[0111] A GPU resource calculation module that calculates the GPU topology distribution of all Pod nodes under the target service when the target service is scaled down;
[0112] A Pod scale-down module that determines the available GPU resource situation based on the GPU topology distribution and performs Pod scale-down.
[0113] The optimization device for heterogeneous resources in the Pod scale-down scenario based on K8s in this embodiment provides a recovery solution dedicated to GPU resource services, which reduces the resource fragmentation process by preferentially vacating more GPU cards. Specifically, when the Pod scale-down module performs scale-down, the GPU resource calculation module calculates the GPU topology distribution of all PODs under the entire service, and preferentially finds the position that can vacate the most GPU resources and only requires scaling down a smaller number for scale-down. In this way, when the Pod is scaled down, through the optimization method for heterogeneous resources in the Pod scale-down scenario based on K8s in this embodiment, more GPU cards can be vacated preferentially, effectively reducing the degree of GPU resource fragmentation.
[0114] Furthermore, the optimization device for heterogeneous resources in the Pod scale-down scenario based on K8s in this embodiment includes:
[0115] An Rs-webhook module that provides the Rs-webhook component by using the capabilities of webhook in K8s;
[0116] When there is a change in the replicaset resource of K8s, call the corresponding interface of the Rs-webhook component and pass the replicaset resources before and after the change as parameters;
[0117] The Rs-webhook component judges the value of the replicas field in the K8s replicaset resource before and after the change to judge whether the target service is scaled down.
[0118] Based on the K8s cluster, when there are changes in the replicaset resource (when the deploy performs scale-down or update, the deploy controller of K8s will operate on the replicas resource. For scale-up or scale-down, that is, modify the value of the replicas field of the current replicaset. When there are changes, a new replicaset will be generated to replace the old replicaset resource, and the value of the replicas field of the old replicaset will be scaled down, while the value of the replicas field of the new replicaset will be scaled up), the kube-apiserver will call the corresponding interface of the Rs-webhook component and pass the replicaset resources before and after the change as parameters.
[0119] The Rs-webhook component will determine whether the replicaset uses GPU resources and the replicas values before and after the change. If it does not use GPU resources or it is a scale-up, no processing will be done. If it is a scale-down, it will enter the processing flow of the heterogeneous resource optimization method in the Pod scale-down scenario.
[0120] Therefore, the Rs-webhook module in this embodiment determines whether the target service is scaled down by the Rs-webhook component judging the value of the replicas field of the K8s replicaset resource before and after the change, and then determines whether to enter the processing flow of the heterogeneous resource optimization method in the Pod scale-down scenario.
[0121] As an optional implementation manner of this embodiment, for the Rs-webhook module in this embodiment, the Rs-webhook component judging the value of the replicas field of the K8s replicaset resource before and after the change and judging whether the target service is scaled down includes:
[0122] If the value of the replicas field of the replicaset resource before the change is greater than the value of the replicas field of the replicaset resource after the change, it is judged that the target service is scaled down; otherwise, it is judged that the target service is not scaled down.
[0123] For a heterogeneous resource optimization device in the Pod scale-down scenario based on K8s in this embodiment, the Rs-webhook component of the Rs-webhook module judges the value of the replicas field of the K8s replicaset resource before and after the change. If it is a scale-down, it will enter the processing flow of the heterogeneous resource optimization method in the Pod scale-down scenario; if it is not a scale-up, no processing will be done.
[0124] In addition, the Rs-webhook component of the Rs-webhook module in this embodiment also determines whether the replicaset uses GPU resources. If it does not use GPU resources, no processing is performed. If it is a scale-down operation, it enters the processing flow of the heterogeneous resource optimization method in the Pod scale-down scenario.
[0125] Specifically, the GPU resource calculation module in this embodiment calculates the GPU topology distribution of all Pod nodes under the target service, including:
[0126] The Rs-webhook component obtains the number N of Pod nodes to be scaled down, obtains all Pod nodes under the replicas field value of the corresponding pre-change replicaset resources, and the Node nodes corresponding to the Pod nodes.
[0127] For any Node node, obtain the corresponding idle GPU number G1 on this Node node, obtain the number N1 of Pod nodes belonging to the replicaset resource that is currently being scaled down on this Node node, and the GPU resource calculation module calculates the available free GPU number G2 and the scale-down number N2 of this Node node.
[0128] The GPU resource calculation module calculates G2 = G1 + N2 and traverses all Node nodes.
[0129] The GPU resource calculation module obtains the available free GPU number G2 and the scale-down number N2 of all Node nodes and sorts all Pod nodes.
[0130] In this embodiment, the GPU resource calculation module obtains the available free GPU number G2 and the scale-down number N2 of all Node nodes and sorts all Pod nodes. When the Pod scale-down module performs Pod scale-down, it can vacate more GPU cards first, effectively reducing the degree of GPU resource fragmentation.
[0131] As an alternative implementation of this embodiment, for any Node node, obtaining the corresponding idle GPU number G1 on this Node node, obtaining the number N1 of Pod nodes belonging to the replicaset resource that is currently being scaled down on this Node node, and the GPU resource calculation module calculating the available free GPU number G2 and the scale-down number N2 of this Node node includes:
[0132] When the replicaset resource to be scaled down is an exclusive service, the calculation condition is: judge that if N > N1, then let N2 = N1; otherwise let N2 = N.
[0133] As an alternative implementation of this embodiment, for any Node node, obtain the corresponding number of idle GPUs G1 on this Node node, and obtain the number of Pod nodes N1 affiliated with the replicaset resource that is currently being scaled in on this Node node. The GPU resource calculation module calculates the available number of GPUs G2 and the number of nodes N2 that can be scaled in on this Node node, including:
[0134] When the replicaset resource being scaled in is a shared service, the calculation condition is: traverse all GPUs on any Node node, obtain the number of Pod nodes N3 affiliated with the replicaset resource that is currently being scaled in on the GPU, calculate whether the GPU is available G3 and the number of nodes N4 that can be scaled in on the GPU. The calculation condition is as follows: when N > N3, let N4 = N3, otherwise N4 = N;
[0135] The GPU resource calculation module calculates G3 = the number of Pod nodes on the GPU - N4, and stores the relationship between the GPU, pod, and G3 in the memory.
[0136] Furthermore, the GPU resource calculation module of this embodiment obtains the available number of GPUs G2 and the number of nodes N2 that can be scaled in for all Node nodes, and sorts all Pod nodes, including:
[0137] When the replicaset resource being scaled in is an exclusive service, after obtaining the available number of GPUs G2 and the number of nodes N2 that can be scaled in for all Node nodes, sort them in descending order of G2 and ascending order of N2;
[0138] When the replicaset resource being scaled in is a shared service, then obtain the corresponding value of G3 from the memory, filter out the Pod nodes with G3 equal to 0, divide the Pod node list into two groups: 0 and non-0, sort them in descending order of G2 and ascending order of N2 respectively, and merge the result lists into one list in the order of the 0 group and the non-0 group.
[0139] Furthermore, the Pod scaling-in module of this embodiment determines the situation of available GPU resources based on the GPU topology distribution and performs Pod scaling-in, including:
[0140] According to the sorting results of all Pod nodes, select N Pod nodes starting from the beginning, call the kube-apiserver component of K8s to mark pod-deleted-cost as -1, and the process ends and returns successfully.
[0141] An optimization device for heterogeneous resources in the scenario of Pod scaling down based on K8s in this embodiment calculates the priority order of Pods in real time through the Rs-webhook component during the scaling down of replicaset and controls it through pod-deleted-cost; by vacating more resources and reducing fragmentation, the GPU resources are improved, especially the allocation rate in the scenario of mixing multiple resource specifications services.
[0142] Embodiment 2
[0143] The following describes an embodiment of the electronic device of the present invention. This electronic device can be regarded as a specific physical implementation manner of the above method and device embodiments of the present invention. For the details described in the embodiment of the electronic device of the present invention, it should be regarded as a supplement to the above method or device embodiments; for the details not disclosed in the embodiment of the electronic device of the present invention, reference can be made to the above method or device embodiments to implement.
[0144] Figure 3 It is a schematic structural diagram of an electronic device according to an embodiment of the present invention. The electronic device includes a processor and a memory. The memory is used to store computer-executable programs. When the computer program is executed by the processor, the processor executes an optimization method for heterogeneous resources in the scenario of Pod scaling down based on K8s in Embodiment 1.
[0145] As Figure 3 shown, the electronic device is presented in the form of a general computing device. The processor can be one or multiple and work together. The present invention does not exclude distributed processing, that is, the processors can be dispersed in different physical devices. The electronic device of the present invention is not limited to a single entity, but can also be the sum of multiple physical devices.
[0146] The memory stores computer-executable programs, usually machine-readable codes. The computer-readable program can be executed by the processor so that the electronic device can execute the method of the present invention or at least part of the steps in the method.
[0147] The memory includes volatile memory, such as a random access storage unit (RAM) and / or a cache storage unit, and can also be non-volatile memory, such as a read-only storage unit (ROM).
[0148] Optionally, in this embodiment, the electronic device further includes an I / O interface, which is used for the electronic device to exchange data with external devices. The I / O interface can represent one or more of several bus structures, including a memory bus or a memory controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any bus structure in multiple bus structures.
[0149] It should be understood that Figure 3 the electronic device shown is merely an example of the present invention, and the electronic device of the present invention may further include elements or components not shown in the above example. For example, some electronic devices further include a display unit such as a display screen, and some electronic devices further include human-computer interaction elements such as buttons, keyboards, etc. As long as the electronic device can execute the computer-readable program in the memory to implement at least part of the steps of the method of the present invention or the method, it can be considered as the electronic device covered by the present invention.
[0150] Figure 4 is a schematic diagram of a computer-readable recording medium according to an embodiment of the present invention. As Figure 4 shown, a computer-executable program is stored in the computer-readable recording medium. When the computer-executable program is executed, it implements a method for optimizing heterogeneous resources in a Pod scale-down scenario based on K8s according to Embodiment 1 of the present invention. The computer-readable recording medium may include data signals propagated in a baseband or as part of a carrier wave, in which readable program codes are carried. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The readable recording medium may also be any readable medium other than the readable recording medium, and this readable medium may send, propagate, or transmit a program used by or in conjunction with an instruction execution system, apparatus, or device. The program codes contained on the readable recording medium may be transmitted by any appropriate medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination of the above.
[0151] The program codes for performing the operations of the present invention may be written in any combination of one or more programming languages. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program codes may be executed entirely on the user computing device, partially on the user device, executed as an independent software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (for example, by using an Internet service provider to connect through the Internet).
[0152] Through the above description of the embodiments, those skilled in the art can easily understand that the present invention can be implemented by hardware capable of executing a specific computer program, such as the system of the present invention, as well as the electronic processing unit, server, client, mobile phone, control unit, processor, etc. included in the system. The present invention can also be implemented by computer software that executes the method of the present invention, such as control software executed by a microprocessor, electronic control unit, client, server side, etc. However, it should be noted that the computer software that executes the method of the present invention is not limited to being executed in one or specific hardware entities, and it can also be implemented in a distributed manner by unspecified specific hardware. For computer software, the software product can be stored in a computer-readable recording medium (which can be a CD-ROM, USB flash drive, mobile disk, etc.), or can be distributed and stored on the network, as long as it can enable an electronic device to execute the method according to the present invention.
[0153] The above embodiments are only used to illustrate the present invention and do not limit the technical solutions described in the present invention. Although this specification has described the present invention in detail with reference to the above respective embodiments, the present invention is not limited to the above specific embodiments. Therefore, any modification or equivalent replacement to the present invention; and all technical solutions and their improvements that do not depart from the spirit and scope of the invention are covered by the scope of the claims of the present invention.
Claims
1. A heterogeneous resource optimization method based on K8s Pod scaling scenario, characterized in that: include: When the target service is scaled down, the GPU topology distribution of all Pod nodes under the target service is calculated; Based on the GPU topology distribution, the available GPU resources are determined and the Pod is scaled down.
2. According to a method for optimizing heterogeneous resources in a Pod scaling scenario based on K8s according to claim 1, it is characterized in that: include: By using the webhook capability in K8s, the Rs-webhook component is provided; When the replicaset resources of K8s are changed, the corresponding interface of the Rs-webhook component is called, and the replicaset resources before and after the change are passed in as parameters; The Rs-webhook component determines the replicas field value before and after the replicaset resource of K8s is changed, and determines whether the target service is scaled down.
3. According to a method for optimizing heterogeneous resources in a Pod scaling scenario based on K8s according to claim 2, it is characterized in that: The Rs-webhook component determines the replicas field value before and after the replicaset resource of K8s is changed, and determines whether the target service is scaled down, including: If the replicas field value of the replicaset resource before the change is greater than the replicas field value of the replicaset resource after the change, it is determined that the target service is being scaled down; otherwise, it is determined that the target service is not being scaled down.
4. According to a method for optimizing heterogeneous resources in a Pod scaling scenario based on K8s according to claim 2 or 3, it is characterized in that: The GPU topology distribution of all Pod nodes under the computing target service includes: The Rs-webhook component obtains the number N of Pod nodes that have been scaled down, obtains all Pod nodes under the replicas field value of the replicaset resource before the change, and the Node nodes corresponding to the Pod nodes; For any Node node, obtain the number of idle GPUs G1 corresponding to the Node node, obtain the number of Pod nodes N1 on the Node node that belongs to the replicaset resource that is currently being scaled down, and calculate the number of available GPUs G2 and the number of Node nodes that can be scaled down N2; Calculate G2=G1+N2 and traverse all Node nodes; Get the available GPU quantity G2 and the scalable capacity N2 of all Node nodes, and sort all Pod nodes.
5. According to a method for optimizing heterogeneous resources in a Pod scaling scenario based on K8s according to claim 4, it is characterized in that: For any Node node, obtaining the number of idle GPUs G1 corresponding to the Node node, obtaining the number of Pod nodes N1 on the Node node belonging to the replicaset resource currently being scaled down, and calculating the number of available GPUs G2 and the number of scalable Node nodes N2 include: When the scaled-down replicaset resource is an exclusive service, the calculation condition is: if N>N1, then let N2=N1; otherwise, let N2=N.
6. According to a method for optimizing heterogeneous resources in a Pod scaling scenario based on K8s according to claim 4, it is characterized in that: For any Node node, obtaining the number of idle GPUs G1 corresponding to the Node node, obtaining the number of Pod nodes N1 on the Node node belonging to the replicaset resource currently being scaled down, and calculating the number of available GPUs G2 and the number of scalable Node nodes N2 include: When the replicaset resource being scaled down is a shared service, the calculation condition is as follows: traverse all GPUs of any Node node, obtain the number of Pod nodes N3 on the GPU that belong to the replicaset resource currently being scaled down, calculate whether the GPU has free space G3 and the number of Pod nodes N4 that can be scaled down on the GPU, and the calculation condition is as follows: when N>N3, set N4=N3, otherwise N4=N; Calculate G3 = the number of Pod nodes on the GPU – N4, and store the relationship among GPU, pod, and G3 in memory.
7. According to a method for optimizing heterogeneous resources in a Pod scaling scenario based on K8s according to claim 4, it is characterized in that: The method of obtaining the available GPU quantity G2 and the scalable capacity N2 of all Node nodes and sorting all Pod nodes includes: When the replicaset resource to be scaled down is an exclusive service, obtain the available GPU quantity G2 and the scalable quantity N2 of all Node nodes, and sort them in descending order of G2 and ascending order of N2; When the scaled-down replicaset resource is a shared service, the value corresponding to G3 is obtained from the memory, and the Pod nodes with G3 of 0 are filtered out. The Pod node list is divided into two groups, 0 and non-0, and arranged in reverse order of G2 and in ascending order of N2 respectively. The result list is merged into one list in the order of group 0 and group non-0.
8. According to a method for optimizing heterogeneous resources in a Pod scaling scenario based on K8s according to claim 6 or 7, it is characterized in that: Determining the free GPU resources based on the GPU topology distribution and scaling down the Pod includes: According to the sorting results of all Pod nodes, N Pod nodes are selected from the beginning, and the kube-apiserver component of K8s is called to mark the pod-deleted-cost as -1. The process ends and returns success.
9. A heterogeneous resource optimization device in a Pod scaling scenario based on K8s, characterized in that: include: GPU resource calculation module, when the target service is scaled down, calculates the GPU topological distribution of all Pod nodes under the target service; The Pod reduction module determines the free GPU resources based on the GPU topology distribution and reduces the Pod capacity.
10. A computer-readable recording medium storing a computer-executable program, characterized in that: When the computer executable program is executed, a heterogeneous resource optimization method in a Pod scaling scenario based on K8s is implemented as described in any one of claims 1 to 8.