Intelligent horizontal Pod automatic expansion system oriented to micro-service architecture
Through the intelligent level Pod automatic expansion system, Kubernetes HPA's insufficient resource coordination in the microservice architecture is solved, efficient dynamic allocation and optimization of resources is achieved, resource utilization and stability of the microservice architecture is improved, and it is suitable for high-concurrency scenarios such as e-commerce and financial transactions.
Patent Information
- Application Number
- CN202510312534.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-07-04
AI Technical Summary
Traditional Kubernetes HPA cannot achieve resource coordination from a global perspective in microservice architecture, resulting in resource fragmentation and elastic imbalance, especially in high-load scenarios, which may lead to insufficient or over-configuration of key service resources, affecting overall application performance and stability.
Design an intelligent horizontal Pod automatic expansion system, and dynamically adjust resource allocation between microservices through distributed system resource management and adaptive software architecture, combining multi-strategy collaboration and resource efficiency heuristic algorithms to achieve cross-service resource redistribution and optimization.
It significantly improves resource utilization and elastic expansion capabilities, reduces resource waste and insufficient allocation, enhances the reliability of the architecture and decision-making coordination efficiency, and is suitable for the microservice expansion needs in high concurrency scenarios.
Smart Images

Figure CN120256016A_ABST
Abstract
Description
Technical Field
[0001] This invention patent belongs to the technical field of cloud computing and containerization, and specifically relates to an intelligent horizontal Pod auto-scaling system for microservices architecture. Background Art
[0002] With the wide application of microservices architecture in the field of cloud computing, enterprises have significantly improved development agility and system fault tolerance by decomposing complex applications into modular services with independent deployment and lightweight communication (such as order processing, payment gateway, and user authentication modules in an e-commerce system); however, this architecture highly depends on the dynamic resource management capabilities of container orchestration platforms (such as Kubernetes). Especially in scenarios of traffic peaks, due to the inability to break through the single-service resource ceiling and the lack of cross-service resource coordination strategies in traditional auto-scaling mechanisms, high-value business modules (such as spike services) frequently encounter resource exhaustion and downtime, while auxiliary services (such as log collection) are idling due to static quota restrictions for a long time, exposing the core contradictions of resource fragmentation and elastic imbalance.
[0003] Analysis of problems in decentralized architecture: Native Kubernetes HPA makes scaling decisions for each microservice individually. Although it can solve the single-point failure problem of centralized architecture, it lacks a global perspective. In a resource-constrained environment, this may lead to some microservices obtaining excessive resources while other key services are resource-starved, affecting the overall application performance. Due to the lack of coordination mechanism between microservices, the decentralized architecture may lead to frequent scaling operations. Especially in scenarios where interdependent microservices compete for limited resources, this frequent scaling will increase system instability and resource waste. In a resource-constrained environment, native HPA cannot achieve resource exchange between microservices. This means that even if some microservices have surplus resources, they cannot effectively allocate these resources to services in need, resulting in reduced overall resource utilization. Since HPA makes scaling decisions based on the resource capacity of a single microservice and predefined thresholds, in resource-constrained situations, it may cause key services to not obtain sufficient resources (i.e., the expected number of expanded replicas is greater than the maximum replica number threshold), leading to performance degradation or even service unavailability. In this scenario, native Kubernetes HPA only simply instructs to expand the number of replicas to the predefined maximum threshold.
[0004] Analysis of latency issues: The default global scan period of Prometheus is 60 seconds, but its jobs can have their own scan periods. The longer the scraping period, the greater the metric latency; the shorter the scraping period, the greater the pressure on the Prometheus Server. For example, query rate(http_requests_total[1m]) returns the average rate per second of the time series collected within 1 minute. The number of time series depends on the scraping period. The HPA controller can adjust the number of application Pod replicas according to the configuration set by the user in advance, but its internal algorithm logic is a passive scaling strategy. It compares the real-time metric data collected with the thresholds in the configuration file and then determines whether to trigger scaling through a simple algorithm logic. For ordinary microservice applications, the native HPA can handle the scaling operations. However, some microservice applications may face sudden traffic during a specific period or season. For example, an online shopping platform launches a product discount service during holidays, and users will flock to the shopping platform to make purchases at the beginning of the discount period. Some users' shopping requests can be satisfied, while others will face shopping failures. This is because the native HPA only triggers the scaling operation when it observes that the metric data (such as CPU utilization) reaches the threshold. The creation of new Pod replicas involves operations such as Pod scheduling, image pulling, container initialization, and connection to the internal database of the application. Only when these operations are completed will the Pod be considered successfully started, and Kubernetes will forward user request traffic to these newly created Pod replicas. During the entire scaling process mentioned above, the original Pods will be in a high-load state, which may lead to the collapse of the entire service.
[0005] Quantification analysis lag: Using the default Kubernetes resource metrics, the collection period of metrics is the default 60s. The goal is to evaluate the performance of HPA, including CPU usage, the number of Pod replicas, and the failed requests within the default metric collection period. Due to the increase in the request rate, after the average CPU usage rises to the limit, as the number of Pod replicas increases, the CPU usage gradually decreases. It can be noticed that the CPU usage changes in each metric collection period (60 seconds) and remains fixed within 60 seconds. This is because Kubelet only collects the original metrics from cAdvisor at the beginning of the metric collection period and will not collect metrics again within the next 60 seconds until the start of the next period. Due to the increase in CPU utilization from the 35th second to the 40th second, around the 50th second, HPA performed the first scaling-up operation, expanding the number of Pod replicas to 8. As time goes by, the CPU utilization reaches the maximum value of 200%. At the 65th second, the number of Pod replicas is expanded to 15, and the third scaling-up occurs at the 110th second. However, at this time, the CPU utilization is still very high. This is caused by the characteristics of HPA. By default, HPA checks the metric values every 15 seconds. When checking, if the metric values have not changed compared to the metric values checked in the previous period, it is considered that this round of operation is not necessary to trigger scaling. Until the 160th second, the collected CPU usage metrics start to decline, and no more scaling-up operations are triggered. The number of replicas remains stable until the end of the experiment. These requests will fail because the newly created Pods are not ready for the load traffic. Therefore, any requests routed to them during the scaling-up operation will fail. In addition, the first scaling-up will cause most of the requests to fail because it occurs during a high-traffic period, and the high request rate causes more requests to be rejected. Summary of the Invention
[0006] The purpose of the present invention is to propose an intelligent horizontal Pod auto-scaling system for a microservices architecture in view of the deficiencies of the prior art. Through distributed system resource management and adaptive software architecture design, and integrating automated operation and maintenance and intelligent optimization algorithm technologies, it aims to solve the problems of inflexible scaling, resource waste, and low decision-making coordination efficiency of traditional auto-scalers in resource-constrained scenarios.
[0007] The purpose of the present invention is achieved through the following technical solutions: An intelligent horizontal Pod auto-scaling system for a microservices architecture, including the following steps:
[0008] (1) Multiple microservice managers: Each microservice manager is bound to a microservice, used to monitor the resource utilization metrics of the microservice in real time, and generate scaling decisions according to the preset scaling policy;
[0009] (2) Microservice Capacity Analyzer: Connects all microservice managers, used to compare the resource requirements of each microservice with the predefined capacity upper limit. When it detects that the resource requirement of any microservice exceeds its capacity upper limit, it triggers a resource coordination signal;
[0010] (3) Adaptive Resource Manager: Responds to the resource coordination signal, executes a resource efficiency heuristic algorithm, calculates the resource transfer plan between microservices, and dynamically adjusts the replica number and resource capacity upper limit of the target microservice.
[0011] Further, step (3) specifically includes the following steps:
[0012] (3.1) Resource Supply and Demand Classification: For the microservice set S = {s1, s2,..., s M}, define the resource supply surplus set S over and the supply shortage set S under , where:
[0013] S under = {s i ∣ DR i > maxR i}, S over = {s j ∣ DR j < maxR j}
[0014] Where, DR i is the expected replica number of microservice s i , and maxR i is its preset capacity upper limit;
[0015] (3.2) Priority Sorting: Sort S under in descending order according to the resource gap. The resource gap is calculated as:
[0016] ΔR i = DR i - maxR i
[0017] Sort S over in ascending order according to the remaining resources. The remaining resources are calculated as:
[0018] ΔR j = maxR j - DR j
[0019] (3.3) Iterative Resource Transfer: For each s i ∈ S under , extract resources from S over in order until ΔR i is satisfied or the resources are exhausted; In a single iterative step, from sj Transfer to s i The resource amount T j→i is as follows:
[0020] T j→i = min(ΔR j ·ResReq j , ΔR i ·ResReq i )
[0021] where ResReq i is the resource request amount of a single copy of s i ;
[0022] (3.4) Capacity update: Update the capacity upper limits of s i and s j :
[0023]
[0024] where maxR′ i and maxR′ j are the updated capacity upper limits.
[0025] Furthermore, when the remaining resources are insufficient to meet the demand in step (3.3), the resources are allocated proportionally so that the microservice s i with insufficient supply obtains the maximum feasible number of replicas, and the calculation formula is:
[0026]
[0027] At the same time, update the capacity upper limit of the microservice s j with excessive supply:
[0028]
[0029] where FeasibleR i is the maximum feasible number of replicas of the microservice s i .
[0030] Furthermore, the scaling strategy described in step (1) includes:
[0031] (1.1) Threshold-based static strategy: The number of replicas is calculated as
[0032]
[0033] where CR i is the current number of replicas, CMV i is the current metric value, and TMV i is the threshold;
[0034] (1.2) Dynamic Policy Based on Fuzzy Logic: Define a fuzzy rule base, and the membership function under high load is:
[0035]
[0036] (1.3) The state space of the Q-learning model is (u, r, CR), where u is the resource utilization rate, CR is the current number of replicas, r is the request rate, the action space is {ScaleUp, ScaleDown}, and the reward function is:
[0037]
[0038] Among them, TMV represents the threshold of resource utilization rate, and ScaleUp and ScaleDown are the core actions in the action space, representing the decision-making operations of resource expansion and resource contraction respectively.
[0039] Furthermore, the comprehensive score Score of the resource utilization rate index i is calculated as:
[0040]
[0041] where w1 + w2 = 1, CPU i is the CPU resource of the i-th microservice, Latency i is the latency of the i-th microservice, CPU th and Latency th are preset thresholds, and the weights w1, w2 are dynamically adjusted according to the microservice type:
[0042]
[0043] Among them, CPU criticality represents the CPU resource critical value, and Latency criticality represents the latency critical value.
[0044] Furthermore, the adaptive resource manager realizes resource redistribution by modifying the Kubernetes custom resource definition CRD. The specific steps include:
[0045] 1) Define the extended resource descriptor ResourcePool;
[0046] 2) Listen for changes in ResourcePool through the Kubernetes Watch API to trigger replica number adjustment.
[0047] Furthermore, the activation condition for the communication between the microservice manager and the adaptive resource manager to trigger the resource coordination signal is:
[0048] And
[0049] where S = {s1, s2, …, s M} is the microservice set, S over is the microservice resource over - supply set of DR j <maxR j and si represents the i - th microservice therein; DR i is the expected number of replicas of microservice s i , maxR i is the preset capacity upper limit of the expected number of replicas of microservice s i , and ResReq i is the single - replica resource request amount of the j - th microservice in set S j . over
[0050] Furthermore, the system is implemented by extending the Kubernetes HPA controller, and the specific code modifications include:
[0051] a. Insert a hook function CheckCapacity in the HPA controller to call the microservice capacity analyzer;
[0052] b. Rewrite the replica number calculation function calculateReplicas to add resource transfer logic.
[0053] Furthermore, the intelligent horizontal Pod auto - scaling system further includes a knowledge base for storing the real - time resource status, historical scaling decisions, and global resource allocation records of each microservice, providing data support for the resource - efficiency heuristic algorithm.
[0054] Advantages of the present invention:
[0055] 1. Significantly improve resource utilization and elastic scaling ability
[0056] By dynamically allocating redundant resources among microservices through a resource efficiency heuristic algorithm, the intelligent HPA breaks through the static replica upper limit restriction of traditional HPA. Experiments show that in resource-constrained scenarios (such as setting the microservice replica upper limit to 5 in benchmark tests), the intelligent HPA reduces the CPU over-utilization rate by 5 times (from 98.06% of Kubernetes HPA to 19.30%), completely eliminates the under-provisioning phenomenon (the traditional solution has a shortage of 934.04m CPU for 13.46 minutes), and at the same time increases the resource allocation efficiency by 1.8 times (from 6110.41m to 11188.76m CPU resources). This dynamic resource transfer mechanism enables high-load microservices to borrow idle resources from low-load services, avoiding performance bottlenecks caused by fixed quotas and reducing resource fragmentation.
[0057] 2. Enhance Architecture Reliability and Decision Coordination Efficiency
[0058] The hierarchical hybrid architecture combines the advantages of decentralized local decision-making (reducing the risk of single-point failures) and centralized global coordination (avoiding resource contention). Under normal low-load conditions, the decentralized microservice manager runs independently, reducing communication overhead (such as only 1.50 minutes are required to trigger centralized coordination in benchmark tests); when resource conflicts are detected, the adaptive resource manager realizes cross-service resource reallocation through priority sorting (such as processing in descending order of resource shortage degree). Experiments show that it extends the over-provisioning time by 9.74 times (from 1.54 minutes of Kubernetes HPA to 15 minutes), and there are no service interruption events.
[0059] 3. Promote the Standardization of Cloud Native Technologies
[0060] The invention provides a reusable automatic scaling enhancement module (open-sourced implementation) for mainstream platforms such as Kubernetes. Its hierarchical architecture design has been verified to be applicable to high-concurrency scenarios such as e-commerce and financial transactions, providing a reference paradigm for microservice elastic scaling in resource-constrained environments and accelerating the standardization process of cloud native automated operation and maintenance technologies. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0062] Figure 1 It is a schematic diagram of the process implemented by an intelligent horizontal Pod autoscaler system for a microservice architecture provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0063] To make the technical solutions and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.
[0064] As Figure 1 shown, the present invention provides an intelligent horizontal Pod auto-scaling system for a microservices architecture. Through a layered architecture design, it combines decentralized resource monitoring with centralized global optimization, innovatively introduces a resource efficiency heuristic algorithm, allows dynamic sharing of redundant resources among microservices, and breaks through the limitation of the predefined replica upper limit. Specifically, the system deploys an independent microservices manager for each microservice, which collects metrics such as CPU utilization and response time in real time and generates a preliminary scaling decision. When it detects that the resource demand exceeds the capacity, the centralized adaptive resource manager triggers cross-service resource reallocation, preferentially reclaims idle resources from low-load microservices, and allocates them to high-load services according to priorities. Finally, global resource balance is achieved by dynamically adjusting the number of replicas and resource quotas. Compared with traditional solutions, the present invention significantly reduces the over-allocation rate of resources (verified by experiments to be reduced by 7 times) and eliminates under-allocation while ensuring service availability, and is particularly suitable for elastic scaling requirements in high-concurrency scenarios such as e-commerce big promotions and real-time streaming media. The following elaborates in detail from three aspects: core architecture, algorithm design, and technical breakthroughs:
[0065] 1. Hierarchical hybrid architecture design
[0066] 1.1 Decentralized microservices manager (local decision-making layer)
[0067] An independent controller, i.e., a microservices manager, is deployed for each microservice to monitor the resource utilization metrics of the microservice in real time (including CPU usage and response time), and is responsible for monitoring resource metrics in real time and generating a preliminary scaling decision. All microservices managers are connected to a microservice capacity analyzer, which is used to compare the resource demands of each microservice with the predefined capacity upper limit. When it detects that the resource demand of any microservice exceeds its capacity upper limit, a resource coordination signal is triggered. The specific process is as follows:
[0068] Monitoring metric collection: Periodically collect metrics such as CPU utilization u CPU , memory usage u mem , request latency t latency etc., and the sampling frequency F sample is dynamically adjusted according to the service load (e.g., F sample = 10Hz at high load and F sample = 1Hz at low load).
[0069] Extended decision generation: Calculate the expected number of replicas DR of microservice s based on a preset policy i For example, when adopting a threshold-based policy, the calculation formula is: i
[0070]
[0071] where CR i is the current number of replicas, and Threshold CPU is the CPU utilization threshold (default 80%). If DR i > maxR i (preset replica upper limit), it is marked as a service with resource shortage.
[0072] 1.2 Centralized adaptive resource manager (global optimization layer)
[0073] Only activated during resource conflicts, in response to the resource coordination signal, execute a resource efficiency heuristic algorithm, calculate the resource transfer plan between microservices, and dynamically adjust the number of replicas and the resource capacity upper limit of the target microservice; achieve cross-service resource reallocation through a lightweight communication protocol:
[0074] Resource conflict detection: When the expected number of replicas of a certain microservice s i exceeds the preset capacity upper limit maxR i , that is, the communication activation condition between the microservice manager and the adaptive resource manager is:
[0075] and
[0076] where S = {s1, s2,..., s M} is the microservice set, S over is the microservice resource supply surplus set where DR j < maxR j , s i represents the i-th microservice among them; maxR i is the expected number of replicas preset capacity upper limit of microservice s i , and ResReq j is the single-replica resource request volume of the j-th microservice in the set S over .
[0077] Communication mechanism: Adopt the gRPC bidirectional stream protocol, the message format is Protocol Buffers, and define the resource request and response structure:
[0078] message ResourceRequest{
[0079] string service_id = 1; / / Service identifier
[0080] int32 required_replicas = 2; / / Required number of replicas
[0081] double cpu_needed = 3; / / Required CPU resources (unit: millicores)
[0082] }
[0083] message ResourceResponse{
[0084] string donor_service = 1; / / Resource provider service ID
[0085] int32 allocated_replicas = 2; / / Allocated number of replicas
[0086] double cpu_allocated = 3; / / Allocated CPU amount
[0087] }
[0088] The adaptive resource manager realizes resource redistribution by modifying the Kubernetes Custom Resource Definition (CRD). The specific steps include:
[0089] a. Define the extended resource descriptor ResourcePool:
[0090] b. Listen for changes in ResourcePool through the Kubernetes Watch API and trigger replica number adjustment.
[0091] 2. Resource efficiency heuristic algorithm (dynamic resource redistribution)
[0092] 2.1 Resource supply and demand classification and priority ranking
[0093] Set of insufficient supply: Filter all microservices with DR i > maxR i to form the set S of microservices with insufficient resource supply, and sort them in descending order according to the resource gap ΔR under , where i ·ResReq i :
[0094] ΔR i = DR i - maxR i , ResReq i = Resource request amount for a single replica
[0095] Over - supply set: Filter all DR j <maxR j of microservices to form the over - supply set S of microservice resources over , according to the remaining resources ΔR j ·ResReq j sort in ascending order:
[0096] ΔR j = maxR j - DR j
[0097] 2.2 Iterative resource transfer and capacity update
[0098] For each s i ∈S under , transfer resources in order from S over until ΔR i is satisfied or the resources are exhausted:
[0099] Single - transfer amount calculation: The amount of resources T over transferred from s j in S under to s i in S j→i is:
[0100] T j→i = min(ΔR j ·ResReq j , ΔR i ·ResReq i )
[0101] where ResReq i is the resource request amount of a single copy of s i .
[0102] Capacity dynamic adjustment: Update the upper limits of the capacities of s i and s j to maxR′ i and maxR′ j respectively:
[0103]
[0104] For example, if a single copy of s i requires 200m CPU and s j has 300m CPU remaining, then the capacity of s i increases by 1 copy, and the capacity of s j decreases by 1 copy.
[0105] 2.3 Partial resource satisfaction strategy
[0106] When the global remaining resources are not sufficient to fully meet the demand, proportional allocation is carried out so that the microservice s with insufficient supply i obtains the maximum feasible number of replicas, and the calculation formula is:
[0107]
[0108] At the same time, update the capacity limit of the microservice s with excess supply j :
[0109]
[0110] Experiments show that this strategy reduces the service unavailability time from 14.7 minutes per day to 0 in resource-constrained scenarios.
[0111] 3. Multi-strategy collaboration and elasticity metric model
[0112] 3.1 Support for hybrid scaling strategies
[0113] 1) Threshold-based static strategy: The number of replicas is calculated as
[0114]
[0115] where CR i is the current number of replicas, CMV i is the current metric value, and TMV i is the threshold. 2) Fuzzy-logic-based dynamic strategy: Define the membership function of the resource utilization rate u. For example, in the "high load" state:
[0116]
[0117] Combine the rule base (such as "IF μ High (u)>0.7 THEN scale out 2 replicas") to generate a decision.
[0118] 3) Reinforcement learning strategy: Construct a Q-learning model. The state space is (u, CR, r), where u is the resource utilization rate, CR is the current number of replicas, and r is the request rate. The action space is {ScaleUp, ScaleDown}, and the reward function is designed as:
[0119]
[0120] Among them, TMV represents the threshold of resource utilization rate. ScaleUp and ScaleDown are the core actions in the action space, representing the decision-making operations of resource expansion and resource reduction respectively. When the system detects that the resource demand exceeds the current capacity, ScaleUp is triggered to increase the number of Pod replicas; when the system load drops or there is resource redundancy, ScaleDown is triggered to reduce the number of Pod replicas. Experiments show that this strategy reduces over-provisioning by 23% compared to the static threshold strategy in dynamic load scenarios.
[0121] 3.2 Multi-metric weighted scoring model
[0122] Define a comprehensive score Score for critical microservices (such as payment gateways) i , balancing the CPU and latency metrics:
[0123]
[0124] Among them, Latency i represents the latency of the i-th critical microservice, CPU th and Latency th represent the preset CPU resource threshold and the preset latency threshold respectively, and CPU i represents the CPU resource of the i-th critical microservice. The weights w1 + w2 = 1, and w1, w2 are dynamically adjusted according to the service type:
[0125]
[0126] For example, for the payment service, w1 = 0.7 and w2 = 0.3 are set to ensure that CPU resources are prioritized. Among them, CPU criticality represents the CPU resource critical value, and Latency criticality represents the latency critical value.
[0127] The system of the present invention is implemented by extending the Kubernetes HPA controller. The specific code modifications include:
[0128] a. Insert a hook function CheckCapacity into the HPA controller to call the microservice capacity analyzer;
[0129] b. Rewrite the replica number calculation function calculateReplicas to add resource transfer logic:
[0130] func(r *Reconciler) calculateReplicas(hpa *autoscalingv2.HorizontalPodAutoscaler, currentReplicas int32) (int32, error) {
[0131] desiredReplicas := baseHPAAlgorithm(hpa) / / Native algorithm
[0132] if desiredReplicas > hpa.Spec.MaxReplicas {
[0133] / / Trigger intelligent HPA resource coordination
[0134] newMax := GetAdjustedMaxReplicas(hpa.Name)
[0135] hpa.Spec.MaxReplicas = newMax
[0136] }
[0137] return desiredReplicas, nil
[0138] }
[0139] The present invention also sets up a knowledge base for storing the real-time resource status, historical scaling decisions, and global resource allocation records of each microservice, providing data support for the resource efficiency heuristic algorithm. Through a layered architecture, a dynamic resource allocation algorithm, and a multi-strategy collaboration mechanism, the present invention significantly improves the elastic scaling ability of microservice applications. Experimental data verifies its advantages in resource utilization, cost control, and service stability, providing an efficient automated operation and maintenance solution for cloud native scenarios.
[0140] The above embodiments are used to explain the present invention rather than limit the present invention. Any modifications and changes made within the spirit and scope of the claims of the present invention fall within the protection scope of the present invention.
Claims
1. An intelligent horizontal Pod auto-scaling system for a microservices architecture, characterized in that It includes the following steps: (1) Multiple microservice managers: Each microservice manager is bound to a microservice, used to monitor the resource utilization metrics of the microservice in real time, and generate an expansion decision according to a preset expansion strategy; (2) Microservice capacity analyzer: Connects all microservice managers, used to compare the resource requirements of each microservice with a predefined capacity upper limit. When it detects that the resource requirement of any microservice exceeds its capacity upper limit, it triggers a resource coordination signal; (3) Adaptive resource manager: In response to the resource coordination signal, executes a resource efficiency heuristic algorithm, calculates a resource transfer plan between microservices, and dynamically adjusts the number of replicas and the resource capacity upper limit of the target microservice.
2. The system according to claim 1, characterized in that, Step (3) specifically includes the following steps: (3.1) Resource supply and demand classification: For the microservice set S = {s1, s2, …, s M}, define the resource oversupply set S over and the undersupply set S under , where: S under = {s i | DR i > maxR i}, S over = {s j | DR j < maxR j} Among them, DR i is the expected number of replicas of microservice s i , and maxR i is its preset capacity upper limit; (3.2) Priority sorting: For S under Sort in descending order of resource gap, and the resource gap is calculated as: ΔR i = DR i - maxR i For S over Sort in ascending order of remaining resources, where the remaining resources are calculated as: ΔR j = maxR j - DR j (3.3) Iterative resource transfer: For each s i ∈S under , extract resources from S over in order until ΔR i is satisfied or the resources are exhausted; in a single iterative step, the amount of resources T j transferred from s i to s j→i is: T j→i = min(ΔR j ·ResReq j , ΔR i ·ResReq i ) Among them, ResReq i is s i The resource request amount for a single copy; (3.4) Capacity update: Update s i and s j 's upper capacity limit: where maxR ′ i and maxR j ′ are the updated capacity upper limits.
3. The system according to claim 2, wherein In step (3.3), when the remaining resources are insufficient to meet the demand, allocate resources proportionally so that the under-supplied microservice s i obtains the maximum feasible number of replicas, and the calculation formula is: Update the capacity limit of the oversupply microservices s simultaneously j : Among them, FeasibleR i is the maximum feasible replica number of microservice s i .
4. The system according to claim 1, wherein The expansion strategy described in step (1) includes: (1.1) Threshold-based static strategy: The number of replicas is calculated as where CR i is the current number of replicas, CMV i is the current metric value, TMV i is the threshold value; (1.2) Fuzzy logic-based dynamic strategy: Define a fuzzy rule base, and the membership function under high load is: (1.3) The state space of the Q-learning model is (u, r, CR), where u is the resource utilization rate, CR is the current number of replicas, r is the request rate, the action space is {ScaleUp, ScaleDown}, and the reward function is: Among them, TMV represents the threshold of the resource utilization rate. ScaleUp and ScaleDown are the core actions in the action space, representing the decision-making operations of resource expansion and resource contraction respectively.
5. The system according to claim 1, wherein The comprehensive score Score of the resource utilization rate index i is calculated as: where w1 + w2 = 1, CPU i is the CPU resource of the i-th microservice, Latency i is the latency of the i-th microservice, CPU th and Latency th are preset thresholds, and the weights w1, w2 are dynamically adjusted according to the microservice type: Among them, CPU criticality represents the CPU resource critical value, and Latency criticality represents the latency critical value.
6. The system according to claim 1, characterized in that, The adaptive resource manager realizes resource reallocation by modifying the Kubernetes custom resource definition CRD. The specific steps include: 1) Define an extended resource descriptor ResourcePool; 2) Listen for changes in ResourcePool through the Kubernetes Watch API to trigger replica number adjustment.
7. The system according to claim 1, wherein The activation condition for the communication between the microservice manager and the adaptive resource manager to trigger the resource coordination signal is: Among them, S = {s1, s2, …, s M} is the microservice set, S over is the microservice resource oversupply set of DR j <maxR j The microservice with an oversupply of resources, s i represents the i-th microservice among them; DR i is the expected number of replicas of the microservice s i , maxR i is the preset capacity upper limit of the expected number of replicas of the microservice s i , ResReq j is the single-replica resource request volume of the j-th microservice in the set S over .
8. The system according to claim 1, characterized in that, The system is implemented by extending the Kubernetes HPA controller. The specific code modifications include: a. Insert a hook function CheckCapacity in the HPA controller to call the microservice capacity analyzer; b. Rewrite the replica number calculation function calculateReplicas to add resource transfer logic.
9. The system according to claim 1, wherein The intelligent horizontal Pod auto-scaling system also includes a knowledge base, used to store the real-time resource status, historical expansion decisions, and global resource allocation records of each microservice, providing data support for the resource efficiency heuristic algorithm.