Control resource adaptive allocation method and system for large-scale Kubernetes cluster

By establishing a queuing model and dependency model, collecting multi-dimensional indicator data, allocating CPU resources and predicting loads, the problem of resource contention of management components in large-scale Kubernetes clusters is solved, and resource allocation rationality and cluster performance are improved.

CN120179382APending Publication Date: 2025-06-20ZHEJIANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510161022.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-13
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

Large-scale Kubernetes clusters under high loads have reduced performance, increased request latency, reduced scheduling throughput and frequent component crashes due to resource competition among management components.

Method used

By establishing a queuing model for component request processing and a dependency model between components, multi-dimensional indicator data are periodically collected, idle and busy components are distinguished, minimum CPU resources are allocated to idle components, predict the load of busy components, and optimal resource allocation problems in different states are formed, and optimal solution for adaptive resource allocation is obtained by solving.

Benefits of technology

It effectively improves the rationality of the allocation of management and control resources, alleviates the problem of resource competition, and optimizes the overall performance of the cluster.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120179382A_ABST
    Figure CN120179382A_ABST
Patent Text Reader

Abstract

The invention provides a management and control resource self-adaptive allocation method and system for a large-scale Kubernetes cluster, and the method comprises the steps: periodically collecting multi-dimensional index data, predicting the load of a component, judging the states of a main node and each component, and solving an optimal resource allocation problem in different states, thereby obtaining a resource self-adaptive allocation optimal solution. Through a resource allocation rule and a feedback adjustment mechanism, the legality and effectiveness of an optimal solution of resource adaptive allocation are ensured, the reasonability of management and control resource allocation is improved, the reasonability of management and control resource allocation is effectively improved, the problem of resource contention is relieved, and the overall performance of a cluster is optimized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of adaptive allocation, and particularly relates to a method and system for adaptively allocating control resources for large-scale Kubernetes clusters. Background Art

[0002] In recent years, with the rapid development of cloud computing and container technology, more and more different forms of workloads, including microservices, batch processing tasks, function computing, etc., are deployed on cloud clusters in the form of containers. As the de facto standard for container orchestration, Kubernetes is widely used in cloud clusters. However, the performance and availability of large-scale Kubernetes clusters will both decrease significantly, specifically manifested as increased request latency, decreased scheduling throughput, and frequent crashes caused by component memory overload. The main reason is that under high load, there are serious resource contention problems among multiple control components of the cluster master node: API server, scheduler, controller manager, and etcd.

[0003] Currently, existing Kubernetes cluster resource management and scheduling allocation systems lack resource allocation strategies for control components and cannot effectively solve the resource contention problem; and in clusters with multiple master nodes, the scheduler and controller manager work in a master-slave mode, and the existing strategy randomly places the main instances of the components, which may cause multiple main instances to be concentrated on the same master node, resulting in an exacerbation of the resource contention problem. Most of the existing resource allocation methods are designed specifically for data-plane workloads and cannot meet the requirements of efficient and stable scheduling of control component resources. Summary of the Invention

[0004] Based on the above deficiencies of the existing technology, the present invention provides a method and system for adaptively allocating control resources for large-scale Kubernetes clusters.

[0005] The first aspect of the present invention provides a method for adaptively allocating control resources for large-scale Kubernetes clusters, including the following steps:

[0006] S1: Establish a queuing model for the component request processing process and a dependency model between components.

[0007] S2: Periodically collect multi-dimensional metric data to distinguish idle and busy components.

[0008] S3: Allocate the minimum required CPU resources for idle components and predict the load of busy components.

[0009] S4: Determine the master node and the status of each component according to the load, and formalize the optimal resource allocation problem in different states.

[0010] S5: Solve the optimal resource allocation problem to obtain the optimal solution for resource adaptive allocation.

[0011] S6: Convert the optimal solution into the final resource allocation decision according to the resource allocation rules, and feedback to adjust the resource allocation.

[0012] The second aspect of the present invention provides a control resource adaptive allocation system for a large-scale Kubernetes cluster, including:

[0013] An index collector, which is used to periodically collect multi-dimensional index data and update the mapping relationship between the component load and the CPU occupancy and concurrency of the API server.

[0014] A resource recommender, which is used to distinguish between idle and busy components, allocate the minimum required CPU resources to idle components, predict the load of busy components, and solve the optimal resource allocation problem to obtain the optimal solution for resource adaptive allocation.

[0015] A component updater, which is used to convert the optimal solution into the final resource allocation decision according to the resource allocation rules and feedback to adjust the resource allocation.

[0016] The technical effects produced by the present invention:

[0017] The present invention provides a control resource adaptive allocation method and system for a large-scale Kubernetes cluster. By periodically collecting multi-dimensional index data, predicting the component load, determining the status of the master node and each component, and then solving the optimal resource allocation problem in different states, the optimal solution for resource adaptive allocation is obtained. Through the resource allocation rules and the feedback adjustment mechanism, the legality and effectiveness of the optimal solution for resource adaptive allocation are ensured, the rationality of control resource allocation is improved, the resource contention problem is effectively alleviated, and the overall performance of the cluster is optimized.

[0018] Other features and advantages of the present invention will be described in the following specification, and, in part, will be obvious from the specification, or will be understood by implementing the present invention. The objectives and other advantages of the present invention can be achieved and obtained by the structures specifically pointed out in the written specification, claims, and drawings. Description of the Drawings

[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art, and form a part of the specification, which is used to explain the present invention together with the embodiments of the present invention, and does not constitute a limitation to the present invention. In the drawings:

[0020] Figure 1 This is the state transition diagram corresponding to the queuing model of the component request processing process in the embodiment of the present application;

[0021] Figure 2 This is the flowchart of a method for adaptively allocating management and control resources for a large-scale Kubernetes cluster in the embodiment of the present application;

[0022] Figure 3 This is the structural diagram of a system for adaptively allocating management and control resources for a large-scale Kubernetes cluster in the embodiment of the present application. Detailed implementation manners

[0023] To make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific implementation manners described herein are only used to explain the present invention and do not limit the protection scope of the present invention.

[0024] The embodiment of the present application provides a method for adaptively allocating management and control resources for a large-scale Kubernetes cluster, including the following steps:

[0025] First, establish a queuing model for the component request processing process, classify according to component characteristics and establish a dependency model between components. When performing resource allocation, periodically collect multi-dimensional metric data to distinguish between idle and busy components; allocate the minimum required CPU resources to idle components, predict the load of busy components; judge the status of the master node and each component according to the load, formalize the optimal resource allocation problem in different states; solve the optimal resource allocation problem to obtain the optimal solution for adaptive resource allocation; finally, make a resource allocation decision according to the resource allocation rules and feedback to adjust the resource allocation, effectively improving the rationality of management and control resource allocation, alleviating the resource contention problem, and optimizing the overall performance of the cluster.

[0026] As Figure 1 shown, the embodiment of the present application takes into account the non-linear and dynamic characteristics of the mapping between CPU occupancy and concurrency, accurately models the request process of components, establishes the relationship between cluster performance metrics and resource allocation, lays a foundation for data collection and solving the optimal resource allocation problem, and models the request processing process of management and control components in the Kubernetes cluster as a queuing model, including:

[0027] First, after a request arrives at a component, according to whether the component concurrency reaches the maximum concurrency and whether the component queued request number reaches the maximum queue length, the request may be in an execution, queuing, or rejection state, specifically:

[0028] 1.1) If the component concurrency does not reach the maximum concurrency f *, the request enters the execution state, and during execution, the request generates a CPU load c and a memory load m on the component e .

[0029] 1.2) If the component concurrency reaches the maximum concurrency and the number of queued requests in the component does not reach the maximum queue length q * , the request enters the queuing state, and during queuing, the request generates a memory load m on the component q .

[0030] 1.3) If the component concurrency reaches the maximum concurrency and the number of queued requests in the component reaches the maximum queue length, the component rejects the request.

[0031] The above request processing process can be represented as a queuing process. When there are n requests in the queuing system, the request arrival rate λ n can be expressed as the following formula:

[0032]

[0033] where, Δt represents the cycle length, and n r represents the number of requests processed and completed within the cycle.

[0034] The service rate μ n can be expressed as the following formula:

[0035]

[0036] where, f(x) represents the mapping of CPU occupancy to concurrency, and c * represents the CPU resources allocated to the component, c represents the CPU load, specifically referring to the additional CPU time slices within the cycle.

[0037] The steady-state condition of the queuing model is that the service rate exceeds the request arrival rate, which can be formally expressed as the following formula:

[0038] min(c * , f(f * ))c / n r >n r / Δt

[0039] 2.1) When the queuing system is in a steady state, the state probability distribution of the queuing system can be obtained according to the flow balance equation. Based on the probability distribution, the performance indicators of the queuing system can be further obtained, including the request rejection probability p reject , the average number of requests L, the average number of queued requests L q , the average number of executing requests L e , the average request delay W, the average request queuing delay W q and the average request execution delay W e .

[0040] 2.2) When the queuing system is in a non-steady state, the performance metrics of the queuing system can be estimated, including the average number of queued requests \(L\) q and the average number of executed requests \(L\) e as well as the CPU load \(l\) that cannot be processed in time.

[0041] Whether the queuing system is in a steady state or not, the memory occupancy of the component can be estimated by the following formula:

[0042] \(m_0+(L q m q +L e m e ) / n r

[0043] where \(m_0\) represents the memory occupancy of the component under no load.

[0044] Next, considering the characteristics and interdependencies of different components, directly optimize the performance of components for SLO, accurately model the impact of resource allocation on different components, and lay a foundation for data collection and solving the optimal resource allocation problem. According to the respective service level objectives (Service Level Objective, SLO) of components, the control components are divided into three categories:

[0045] Type A components include API servers and etcd, and their goal is to minimize the request latency \(W\).

[0046] Type B components include schedulers, and their goal is to maximize the scheduling throughput \(S\).

[0047] Type C components include controller managers, and their goal is to timely complete the status coordination of objects in the work queue, that is, to make the queuing system in a stable state. If the CPU resources required to make the queuing system in a steady state is \(x\), then the CPU resources to be allocated to type C components are:

[0048]

[0049] where \(w C is the weight coefficient of type C components, and \(\overline{w}\) is the average value of the weight coefficients of all types of components.

[0050] In particular, among all components, except for type B components that work in an asynchronous mode, the rest of the components work in a synchronous mode.

[0051] Furthermore, there are dependencies between components. Since subsequent requests to the API server and etcd will be generated after the scheduler and controller manager have finished processing, allocating more CPU resources to Class B and Class C components will not only increase their own throughput but also increase the load on Class A components. If the original load ratio of Class B components to Class A components is r B , and the current throughput becomes m times the original, and the original load ratio of Class C components to Class A components is r C , and the current throughput becomes n times the original, then the load on Class A components becomes 1+(m - 1)r B +(n - 1)r C times.

[0052] As Figure 2 shown, an embodiment of this application provides a method for adaptively allocating management and control resources for a large-scale Kubernetes cluster, including:

[0053] Step 1: Periodically collect multi-dimensional metric data, distinguish idle and busy components based on the data, and collect multi-dimensional, differentiated, and easily obtainable metric data for different components according to the queuing model, component classification, and the dependency model between components, accurately describe the component load, distinguish idle and busy components, and lay a foundation for subsequent load prediction and solving the optimal resource allocation problem.

[0054] Furthermore, the specific content of Step 1 is as follows:

[0055] Step 1.1: The general metrics collected include the incremental metrics of the component in the previous cycle and the current latest status metrics.

[0056] The incremental metrics specifically include the number of requests n r processed, the CPU time slice c, the allocated memory size m, and the size s of the network data packets received (which can approximately represent the memory load generated when requests are queued).

[0057] The status metrics specifically include the number of queued requests n q , the CPU occupancy c0 and memory occupancy m0 when the component is idle, and the CPU throttling percentage τ.

[0058] At the same time, special metrics are additionally collected for different components based on component classification:

[0059] For the scheduler and controller manager, an additional metric indicating whether the component is the main instance currently is collected.

[0060] For the API server, the number of requests n r,B from Class B components and n r,C, and considering that the mapping between CPU occupancy and the number of concurrences is non-linear and dynamically changing, additionally collect the data points of (number of concurrences, CPU occupancy) in the previous cycle to update the mapping between CPU occupancy and the number of concurrences.

[0061] Step 1.2: Distinguish idle and busy components based on the collected metrics; for the API server and etcd, the component is idle if and only if n r = n q = 0; for the scheduler and the controller manager, the component is idle if and only if n r = n q = 0, or the component is not the primary instance.

[0062] Step 2: Allocate the minimum required CPU resources for idle components, predict the load of busy components in the next cycle. By allocating the minimum required CPU resources for idle components, as many CPU resources as possible are reserved for busy components, which helps improve the overall performance of the cluster; by predicting the load of busy components in the next cycle, it lays a foundation for the formalization and solution of the subsequent optimal resource allocation problem.

[0063] Furthermore, the specific content of Step 2 is as follows:

[0064] Step 2.1: Allocate CPU resources with a size of max(c0, c req ) to idle components, where c req represents the value of the resources.requests.cpu field in the component configuration file.

[0065] Step 2.2: For busy components, according to the collected incremental metric data, apply the time series prediction algorithm to predict the load in the next cycle, specifically including the number of processed completed requests n′ r , CPU load c′, memory load m e ′ during request execution, and memory load m q ′ during request queuing.

[0066] Step 3: Judge the status of the master node according to the load, formalize the optimal resource allocation problem under different statuses, accurately judge the status of the master node and each component according to the component load, discuss in detail the optimization objectives and constraints under different statuses, accurately formalize the optimal resource allocation problem, and give the optimal allocation strategies in some simple statuses to simplify the subsequent problem-solving process.

[0067] Furthermore, the specific content of Step 3 is as follows:

[0068] Step 3.1: Judge whether the master node is in a contention state according to whether the sum of the predicted loads of busy components in the next cycle exceeds the threshold, specifically as follows:

[0069]

[0070] Among them, Δt represents the cycle length, β represents the resource contention threshold coefficient, and C represents the available CPU resources of the master node.

[0071] 1) If the master node is in a non-contention state, no resource allocation is required, and it waits for the next cycle.

[0072] 2) If the master node is in a contention state, reduce the throughput rate of type B components to 0, adjust the component load according to the component dependency model, and then determine whether the master node is in a globally stable state based on whether the sum of the adjusted component CPU loads exceeds the allocable CPU resources.

[0073] Step 3.2: Regardless of whether the master node is in a globally stable state, the unified constraint conditions for the optimal resource allocation problem are denoted as as follows:

[0074] First, the total sum of the allocated CPU resources does not exceed the available CPU resources of the master node;

[0075] Second, the request rejection probability does not exceed a pre-set threshold P.

[0076] Step 3.3: Formalize the optimal resource allocation problem under different states of the master node.

[0077] 1) If the master node is in a non-globally stable state, the optimization goal is to minimize the total CPU load that all components fail to process in time and balance the time required for each component to process this load. The optimization problem is formalized as the following formula:

[0078]

[0079] Among them, l represents the CPU load that the component fails to process in time, t represents the time required to process this load, and c * represents the CPU resources allocated to the component.

[0080] 2) If the master node is in a globally stable state, determine whether the busy components working in synchronous mode on the master node are in a locally stable state according to whether the following formula holds:

[0081]

[0082] 2.1) For components in a non-locally stable state, allocate the minimum CPU resources required to achieve the steady-state condition

[0083] 2.2) For components in a locally stable state, the optimization goal is to minimize the average request latency of type A components and maximize the scheduling throughput rate of type B components.

[0084] The optimization problem is formulated as follows:

[0085]

[0086] where, represents the average request delay, and w A and w B represent the weight coefficients.

[0087] Step 4: Solve the optimal resource allocation problem to obtain the optimal solution for resource adaptive allocation. Control the solution space size through the CPU granularity adaptive mechanism and the heuristic algorithm to improve the problem-solving efficiency; by solving the optimal resource allocation problems in different states, obtain the optimal solution for resource adaptive allocation, effectively improve the rationality of managed resource allocation, and optimize the overall performance of the cluster.

[0088] Furthermore, the specific steps of Step 4 are as follows:

[0089] 1) For the optimal resource allocation problem when the master node is in a non-globally stable state, an analytical solution c * = c′C / ∑c′ can be obtained, that is, the CPU allocation is proportional to the predicted CPU load of each component.

[0090] 2) For the optimal resource allocation problem of the components in a locally stable state when the master node is in a globally stable state, a genetic algorithm with CPU granularity adaptation is used for solution, including:

[0091] Step 21: Initialize the population size and the maximum number of iterations in the genetic algorithm, set the population iteration end condition, and establish a chromosome encoding and decoding scheme.

[0092] Step 22: Initialize the population and calculate the fitness value of each chromosome in the current population according to the optimization objective function of the optimization problem.

[0093] Step 23: Use the selection operator, crossover operator, and mutation operator to search for population inheritance respectively, and determine whether the current population meets the population iteration end condition. If it meets, select the chromosome with the highest fitness in the population as the optimal resource allocation; if it does not meet, return to Step 22.

[0094] The calculation of CPU granularity follows the following formula:

[0095] γT / C r

[0096] where, γ represents the adjustment coefficient, T represents the time upper limit for solving the genetic algorithm, and C r represents the remaining available CPU resources of the master node except for the CPU required to satisfy all components in a locally stable state to reach a stable state.

[0097] Determine the traffic control parameters f using a heuristic algorithm * and q * Reduce the solution space. When the CPU allocation c * is determined, take the smallest solution of f(x) = c * as the maximum concurrency f * . If the master node is in a globally stable state, for components in a locally stable state, take the solution of 1 - p reject (c * , f * , q * ) = P as the maximum queue length q * . If the master node is in a globally unstable state, or for components in a locally unstable state when the master node is in a globally stable state, take the queue length as the maximum queue length q * .

[0098] Step 5: Convert the optimal solution into the final resource allocation decision according to the resource allocation rules, and feedback and adjust the resource allocation. Achieve CPU resource allocation and update of traffic control parameters without restarting components to ensure the stability of the cluster; ensure the compliance and consistency of resource updates by formulating resource allocation update rules; timely correct model errors through the feedback adjustment mechanism to improve the effectiveness of resource allocation.

[0099] Achieve the update of CPU resource allocation by modifying the value of the resources.limits.cpu field in the component configuration file without restarting the component; achieve the update of traffic control parameters of the API server component through the APF feature of Kubernetes without restarting the component.

[0100] Formulate resource allocation update rules, including:

[0101] Rule 1: Meet the requirements of the Kubernetes API. For example, the value of the resources.limits.cpu field must be greater than or equal to the value of the resources.requests.cpu field, and the maximum concurrency and the maximum queue length must be positive numbers.

[0102] Rule 2: Keep the resource allocation relatively stable before and after. That is, for the same component of the same master node, the difference in resource allocation between two consecutive cycles cannot exceed vC, where v represents the fluctuation coefficient.

[0103] Furthermore, for the optimal solution that does not meet Rule 1, reject the resource allocation update; for the optimal solution that does not meet Rule 2, if the allocated resources are too few, update the resource allocation to the minimum value within the fluctuation range; if the allocated resources are too large, update the resource allocation to the maximum value within the fluctuation range.

[0104] Based on the CPU throttling percentage of each component, perform negative feedback adjustment on resource allocation and weights between every two allocation cycles, including:

[0105] Mark the components with CPU throttling percentage exceeding the highest threshold as bottleneck components, and mark the components with CPU throttling percentage lower than the lowest threshold as over-allocated components;

[0106] If there are both bottleneck components and over-allocated components at the same time, the over-allocated components immediately allocate a part of their allocated CPU resources to the bottleneck components with the throttling percentage of the bottleneck components as the weight, and this resource allocation is not restricted by Resource Allocation Rule 2;

[0107] If there are bottleneck components, increase the resource allocation weight of the bottleneck components in subsequent cycles.

[0108] As Figure 3 shown, the embodiment of the present application also provides a control resource adaptive allocation system for a large-scale Kubernetes cluster, specifically including:

[0109] An index collector, which is used to periodically collect multi-dimensional index data and update the mapping relationship between component load and CPU occupancy and concurrency of the API server.

[0110] A resource recommender, which is used to distinguish between idle and busy components, allocate the minimum required CPU resources to idle components, predict the load of busy components and solve the optimal resource allocation problem to obtain the optimal solution of resource adaptive allocation.

[0111] A component updater, which is used to convert the optimal solution into a final resource allocation decision according to the resource allocation rules and feedback-adjust the resource allocation.

[0112] The specific embodiments described above have detailed the technical solutions and beneficial effects of the present invention. It should be understood that the above are only the most preferred embodiments of the present invention and are not used to limit the present invention. Any modifications, supplements, equivalent replacements, etc. made within the scope of the principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A method for adaptively allocating management and control resources for large-scale Kubernetes clusters, characterized in that: The following steps are involved: S1: Establish a queuing model for component request processing and a dependency model between components; S2: Periodically collect multi-dimensional indicator data to distinguish between idle and busy components; S3: Allocate the minimum CPU resources required by idle components and predict the load of busy components; S4: Determine the status of the master node and each component based on the load, and formalize the optimal resource allocation problem under different states; S5: Solve the optimal resource allocation problem and obtain the optimal solution for adaptive resource allocation; S6: Convert the optimal solution into the final resource allocation decision according to the resource allocation rules, and provide feedback to adjust the resource allocation.

2. According to claim 1, a method for adaptively allocating management and control resources for a large-scale Kubernetes cluster is characterized in that: The S1 specifically includes: S1.1: After the request arrives at the component, the request is judged to be in the execution, queue or rejection state according to whether the component concurrency reaches the maximum concurrency and whether the number of queued requests of the component reaches the maximum queue length, and the request processing process is modeled as a queue model; S1.2: Classify the management and control components according to their respective service level objectives. There are dependencies between the components, and a dependency model is established based on this.

3. According to claim 2, a method for adaptively allocating management and control resources for a large-scale Kubernetes cluster is characterized in that: S1.2 divides the control components into three categories according to their respective service level objectives: Class A components include API servers and etcd, and their goal is to minimize request latency; Class B components include schedulers, whose goal is to maximize scheduling throughput; Class C components contain the controller manager, whose goal is to complete the state coordination of objects in the work queue in a timely manner, that is, to make the queuing system in a stable state.

4. According to claim 1, a method for adaptively allocating management and control resources for a large-scale Kubernetes cluster is characterized in that: The S3 includes: S3.1: Allocate the minimum CPU resources required by the idle component. The specific value is determined by the CPU usage of the component when it is unloaded and the resource lower limit field in the component configuration file. S3.2: For busy components, predict the load of the next cycle based on the collected indicator data.

5. According to claim 1, a method for adaptively allocating management and control resources for a large-scale Kubernetes cluster is characterized in that: The S4 includes: S4.1: judging whether the master node is in a contention state according to whether the sum of the next cycle predicted loads of the busy components exceeds a threshold; S4.2: Set uniform constraints for the optimal resource allocation problem; S4.3: Formalize the optimal resource allocation problem when the master node is in different states.

6. According to claim 5, a method for adaptively allocating management and control resources for a large-scale Kubernetes cluster is characterized in that: The S4.3 is specifically: If the master node is in a non-globally stable state, the optimization goal is to minimize the total CPU load that all components cannot process in time, and balance the time required for each component to process these loads; If the master node is in a globally stable state, determine whether the busy component working in synchronous mode on the master node is in a locally stable state based on whether the maximum CPU usage of the component exceeds the minimum CPU resource required for the queuing system to reach a steady state; For components in a non-local stable state, allocate the minimum CPU resources required to achieve the steady-state condition. For components in a local stable state, the optimization goal is to minimize the average request delay of class A components and maximize the scheduling throughput of class B components, and the weight coefficient is used to merge the two goals into the same goal.

7. A method for adaptively allocating management and control resources for a large-scale Kubernetes cluster according to claim 1 or 6, characterized in that: The S5 includes: 1) For the optimal resource allocation problem when the master node is in a non-globally stable state, an analytical solution is obtained, that is, the CPU allocation is proportional to the CPU load predicted by each component; 2) For the optimal resource allocation problem of components in a local stable state when the master node is in a global stable state, a CPU-granularity adaptive genetic algorithm is used to solve it.

8. According to claim 7, a method for adaptively allocating management and control resources for a large-scale Kubernetes cluster is characterized in that: The genetic algorithm solution using CPU granularity adaptation in 2) specifically includes: Step 21: Initialize the population size and maximum number of iterations in the genetic algorithm, set the population iteration end condition, and establish the chromosome encoding and decoding scheme; Step 22: Initialize the population and calculate the fitness value of each chromosome of the current population according to the optimization objective function of the optimization problem; Step 23: Use the selection operator, crossover operator and mutation operator to search the population genetics respectively to determine whether the current population meets the population iteration end condition. If so, select the chromosome with the highest fitness in the population as the optimal resource allocation. If not, return to step 22. The CPU granularity is positively correlated with the remaining allocatable CPU resources of the master node, in addition to the CPU required for all components in local stable states to achieve a stable state, so as to limit the size of the solution space; a heuristic algorithm is used to determine the flow control parameters to narrow the solution space.

9. The method for adaptively allocating management and control resources for a large-scale Kubernetes cluster according to claim 1, characterized in that: The S6 includes: Initiate resource allocation update requests to the cluster through the Kubernetes API; formulate resource allocation update rules; and perform negative feedback adjustments on resource allocation and weights based on the CPU current limit percentage of each component between every two allocation cycles.

10. A control resource adaptive allocation system for large-scale Kubernetes clusters, characterized in that: include: The indicator collector is used to periodically collect multi-dimensional indicator data, update component load, and the mapping relationship between the CPU usage and concurrency of the API server; Resource recommender, which is used to distinguish between idle and busy components, allocate the minimum required CPU resources to idle components, predict the load of busy components and solve the optimal resource allocation problem to obtain the optimal solution for adaptive resource allocation; The component updater is used to convert the optimal solution into the final resource allocation decision according to the resource allocation rules and to provide feedback to adjust the resource allocation.