Resource dynamic scaling and quota limiting method and device of cloud load balancer, electronic equipment and storage medium

By directly obtaining business data and resource usage and dynamically adjusting the maximum number of replicas, the adapter dependency issue in the HPA controller is resolved, seamless docking and stable scaling are achieved, maintenance costs are reduced, and business performance is improved.

CN120675995AActive Publication Date: 2025-09-19BEIJING KINGSOFT CLOUD NETWORK TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510970237.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-14
Publication Date
2025-09-19
Estimated Expiration
2045-07-14

AI Technical Summary

Technical Problem

Existing containerized elastic scaling solutions rely on the HPA controller, which requires the development of independent adapters for each new business indicator, increasing deployment configuration and version compatibility maintenance costs. The process of converting business indicators into resource request values ​​is complex, leading to architectural redundancy.

Method used

By regularly polling to obtain business data and resource usage, the maximum number of replicas is dynamically adjusted using the first and second weight indicators, directly participating in the update of HPA objects, skipping the intermediate adaptation link, and achieving seamless integration of business indicators and elastic scaling systems.

Benefits of technology

It eliminates adapter development and maintenance costs, avoids architectural redundancy, reduces deployment configuration and version compatibility maintenance costs, and improves scaling stability and business performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120675995A_ABST
    Figure CN120675995A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a resource dynamic scaling and quota limiting method and device for a cloud load balancer, electronic equipment and a storage medium, and the method comprises the steps: carrying out the regular polling to obtain service data and a resource utilization rate for the current calculation of the maximum copy number of Pod, the service data comprises first index data of the first weight index and second index data of the second weight index; determining a basic maximum copy number based on the first index data and the resource utilization rate; according to the second index data, dynamically adjusting the basic maximum copy number through a hierarchical threshold triggering mechanism to determine the final maximum copy number; and updating the maximum number of the HPA copies in the HPA object by using the final maximum number of the copies. According to the scheme, the final maximum copy number is calculated based on the data of the service indexes, the maximum copy number in the HPA object can be dynamically updated, the whole process completely skips an intermediate adaptation link of converting the service indexes into HPA native indexes, the development and maintenance cost of an adapter is eliminated, and architecture redundancy can be avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of Internet technology, and in particular to a method, device, electronic device, and storage medium for dynamic resource scaling and quota restriction of a cloud load balancer. Background Art

[0002] The current mainstream containerized elastic scaling solutions mainly rely on the HPA (Horizontal Pod Autoscaler) controller to trigger expansion and contraction based on single-dimensional indicators.

[0003] Because the HPA controller supports only basic resource metrics like CPU (Central Processing Unit) and memory by default, business metrics like QPS (Queries Per Second), CPS (Connections Per Second), and ACC (Active Concurrent Connections) must be converted into resource request values ​​through an adapter. This tightly coupled design necessitates the development of a separate adapter for each new business metric, creating architectural redundancy and significantly increasing deployment, configuration, and version compatibility maintenance costs. Summary of the Invention

[0004] In view of this, in order to effectively alleviate the above technical problems, embodiments of the present application provide a method, apparatus, electronic device, and storage medium for dynamic resource scaling and quota restriction of a cloud load balancer.

[0005] In a first aspect, an embodiment of the present application provides a method for dynamically scaling and limiting resource quotas of a cloud load balancer, the method comprising:

[0006] Regularly polling to obtain business data and resource usage used to calculate the current maximum number of Pod replicas, the business data including first indicator data of a first weight indicator and second indicator data of a second weight indicator; wherein the first weight indicator and the second weight indicator are both business indicators that do not require adaptation to the HPA controller;

[0007] Determine a basic maximum number of replicas based on the first indicator data and the resource usage rate;

[0008] According to the second indicator data, dynamically adjusting the basic maximum number of replicas through a hierarchical threshold trigger mechanism to determine a final maximum number of replicas;

[0009] The HPA maximum number of copies in the HPA object is updated using the final maximum number of copies.

[0010] Optionally, as in the aforementioned method, the periodic polling to obtain the business data and resource usage used to calculate the current maximum number of Pod replicas includes:

[0011] Obtain a first resource usage sequence currently collected using a time window;

[0012] Extracting the start and end timestamps of the time window;

[0013] Determining whether the time span of the start and end timestamps is less than a preset time span;

[0014] If it is determined that the time span of the start and end timestamps is less than the preset time span, querying the cluster monitoring system based on the preset time span to obtain the current first data sequence;

[0015] In a case where it is determined that the time span of the start and end timestamps is equal to or greater than the preset time span, querying from the cluster monitoring system based on the time span of the start and end timestamps to obtain the current first data sequence;

[0016] Time-aligning the first resource usage sequence and the first data sequence to determine a second resource usage sequence and the second data sequence;

[0017] The resource usage rate is acquired based on the second resource usage rate sequence, and the first indicator data of the first weight indicator and the second indicator data of the second weight indicator are acquired based on the second data sequence.

[0018] Optionally, as in the aforementioned method, the resource utilization includes the current number of Pods, the current memory utilization, and the current CPU utilization;

[0019] The acquiring the resource usage rate based on the second resource usage rate sequence includes:

[0020] Get the latest number of Pods from the second resource usage sequence;

[0021] Determine the average number of Pods and the Pod standard deviation based on the second resource usage sequence;

[0022] Determine whether a 2σ standard deviation noise filtering mechanism is satisfied based on the latest Pod number, the average Pod number, and the Pod standard deviation;

[0023] If it is determined that the noise filtering mechanism based on the 2σ standard deviation is satisfied, the latest Pod number is determined as the current Pod number;

[0024] If it is determined that the noise filtering mechanism based on the 2σ standard deviation is not satisfied, the latest Pod number at the previous running time is determined as the current Pod number;

[0025] Acquire the latest memory utilization from the second resource utilization sequence, and determine the latest memory utilization as the current memory utilization;

[0026] The latest CPU utilization is obtained from the CPU utilization sequence, and the latest CPU utilization is determined as the current CPU utilization.

[0027] Optionally, as in the aforementioned method, first indicator data of the first weighted indicator and second indicator data of the second weighted indicator are obtained based on the second data sequence;

[0028] Acquire the latest first indicator data from the second data sequence, and determine the latest first indicator data as the first indicator data of the first weighted indicator;

[0029] The latest second indicator data is obtained from the second data sequence, and the latest second indicator data is determined as the second indicator data of the second weight indicator.

[0030] Optionally, as in the aforementioned method, determining the basic maximum number of replicas based on the first indicator data and the resource usage rate includes:

[0031] Obtaining a maximum indicator threshold, a memory utilization threshold, and a CPU utilization threshold of the first weighted indicator;

[0032] Determine the maximum number of CPU replicas based on the current number of Pods, the maximum indicator threshold, the first indicator data, the CPU utilization threshold, and the current CPU utilization;

[0033] Determine the maximum number of memory replicas based on the current number of Pods, the maximum indicator threshold, the first indicator data, the memory utilization threshold, and the current memory utilization;

[0034] The maximum value is selected from the maximum number of CPU copies and the maximum number of memory copies, and the maximum value is rounded up to obtain the basic maximum number of copies.

[0035] Optionally, as in the aforementioned method, dynamically adjusting the basic maximum number of replicas through a hierarchical threshold trigger mechanism based on the second indicator data to determine the final maximum number of replicas includes:

[0036] Determine a per-core unit value based on the second indicator data and the current number of Pods;

[0037] Detecting whether the per-core unit value exceeds a preset indicator data threshold range; wherein the preset indicator data threshold range is a data range consisting of a preset pressure value and a limit value corresponding to each core;

[0038] When it is detected that the per-core unit value is within the preset indicator data threshold range, obtaining a real-time change rate of the first weight indicator;

[0039] detecting whether the real-time change rate exceeds a preset change rate threshold;

[0040] When it is detected that the real-time change rate exceeds the preset change rate threshold, the basic maximum number of replicas is adjusted according to the preset adjustment number of Pods, and the adjusted basic maximum number of replicas is determined as the final maximum number of replicas;

[0041] In the case where it is detected that the per-core unit value exceeds the preset indicator data threshold range, modifying the basic maximum number of replicas based on the second indicator data, and determining the basic maximum number of replicas after the modification as the final maximum number of replicas;

[0042] When it is detected that the per-core unit value is lower than the preset indicator data threshold range, or when it is detected that the real-time change rate does not exceed the preset change rate threshold, the basic maximum number of copies is determined as the final maximum number of copies.

[0043] Optionally, as in the aforementioned method, modifying the basic maximum number of replicas based on the second indicator data includes:

[0044] determining a desired number of cores based on the per-core unit value and the limit value;

[0045] Determine the number of replicas to be added based on the required number of cores;

[0046] The basic maximum number of replicas is modified based on the number of replica increases.

[0047] Optionally, as in the aforementioned method, before using the final maximum number of copies to update the HPA maximum number of copies in the HPA object, the method further includes:

[0048] Determine a scaling flag value corresponding to the final maximum number of replicas based on the first indicator data, the current number of Pods, the current CPU utilization, the current memory utilization, the maximum indicator threshold, the CPU utilization threshold, the memory utilization threshold, and the maximum number of HPA replicas;

[0049] The final maximum number of replicas and the corresponding scaling flag value are stored in a recommendation array. The recommendation array also includes the final maximum number of replicas determined at multiple historical moments and the corresponding scaling flag values ​​for each of the final maximum numbers of replicas.

[0050] Optionally, as in the aforementioned method, determining the scaling flag value corresponding to the final maximum number of replicas based on the first indicator data, the current number of Pods, the current CPU utilization, the current memory utilization, the maximum indicator threshold, the CPU utilization threshold, the memory utilization threshold, and the maximum number of HPA replicas includes:

[0051] Determine a single Pod CPU utilization based on the first indicator data, the current number of Pods, and the current CPU utilization; and determine a single Pod memory utilization based on the first indicator data, the current number of Pods, and the current memory utilization;

[0052] Determine a memory Pod incremental demand value based on the single Pod memory utilization, the first indicator data, and the maximum indicator threshold; and determine a CPU Pod incremental demand value based on the single Pod CPU utilization, the first indicator data, and the maximum indicator threshold;

[0053] Determine the remaining available Pod capacity value of the CPU based on the CPU utilization threshold, the current CPU utilization, the current number of Pods, and the maximum number of HPA replicas; and determine the remaining available Pod capacity value of the memory based on the memory utilization threshold, the current memory utilization, the current number of Pods, and the maximum number of HPA replicas;

[0054] The scaling flag value corresponding to the final maximum number of replicas is determined based on the final maximum number of replicas, the maximum number of HPA replicas, the memory Pod incremental requirement value, the CPU Pod incremental requirement value, the CPU remaining available Pod capacity value, the memory remaining available Pod capacity value, the first indicator data, the current CPU utilization rate, the current number of Pods, and the latest number of Pods.

[0055] Optionally, as in the aforementioned method, the method further includes:

[0056] The single Pod CPU utilization and the maximum number of CPU replicas are respectively used as smoothing processing objects, and the following operations are performed on the smoothing processing objects:

[0057] Get multiple historical data values ​​and current data values;

[0058] determining a mean and a standard deviation based on the plurality of historical data values ​​and the current data value;

[0059] Determining whether a standard deviation-based noise filtering mechanism is satisfied based on the current data value, the mean, and the standard deviation;

[0060] If it is determined that the standard deviation noise filtering mechanism is satisfied, retaining the current data value;

[0061] If it is determined that the standard deviation-based noise filtering mechanism is not satisfied, the historical data value at the previous moment is retained.

[0062] Optionally, as in the aforementioned method, determining the scaling flag value corresponding to the final maximum number of replicas based on the final maximum number of replicas, the maximum number of HPA replicas, the memory Pod incremental requirement value, the CPU Pod incremental requirement value, the CPU remaining available Pod capacity value, the memory remaining available Pod capacity value, the first indicator data, the current CPU utilization, the current number of Pods, and the latest number of Pods includes:

[0063] Check whether the final maximum number of replicas is the same as the latest number of Pods;

[0064] When it is detected that the final maximum number of replicas is the same as the latest number of Pods, detect whether the final maximum number of replicas is selected from the maximum number of CPU replicas;

[0065] When it is detected that the final maximum number of replicas is selected from the maximum number of CPU replicas, determining whether the first condition or the second condition is met based on the CPUPod increment requirement value and the CPU remaining available Pod capacity value;

[0066] When the first condition or the second condition is met, setting the value of the telescopic flag to a first value;

[0067] When it is detected that the final maximum number of replicas is selected from the maximum number of memory replicas, determining whether the third condition or the fourth condition is met based on the memory Pod incremental demand value and the memory remaining available Pod capacity value;

[0068] When the third condition or the fourth condition is met, setting the value of the telescopic flag to the first value;

[0069] When it is detected that the final maximum number of replicas is different from the latest number of Pods, determining the average data of the first indicator based on the second data sequence;

[0070] Determining whether the fifth condition is satisfied based on the first indicator data and the first indicator average data, or determining whether the sixth condition is satisfied based on the current number of Pods and the current number of Pods at the previous running time;

[0071] When the fifth condition or the sixth condition is met, setting the value of the telescopic flag to the second value;

[0072] If the fifth condition and the sixth condition are not met, detecting whether the current CPU utilization is less than a preset pressure threshold;

[0073] When detecting that the current CPU utilization is less than the preset pressure threshold, setting the scaling flag value to the third value;

[0074] When it is detected that the current CPU utilization is not less than the preset pressure threshold, detecting whether the final maximum number of replicas is less than the latest number of Pods;

[0075] When it is detected that the final maximum number of replicas is less than the latest number of Pods, the scaling flag value is set to the fourth value.

[0076] When it is detected that the final maximum number of replicas is greater than the latest number of Pods, detecting whether the first indicator data is less than a preset initial threshold value;

[0077] When detecting that the first indicator data is less than the preset initial threshold, setting the scaling flag value to the fifth value;

[0078] When detecting that the first indicator data is not less than the preset initial threshold value, detecting whether the final maximum number of replicas is greater than the HPA maximum number of replicas;

[0079] If it is detected that the final maximum number of replicas is not greater than the HPA maximum number of replicas, setting the scaling flag value to the sixth value;

[0080] When it is detected that the final maximum number of replicas is greater than the HPA maximum number of replicas, the scaling flag value is set to the seventh value.

[0081] Optionally, as in the aforementioned method, updating the maximum number of HPA copies in the HPA object using the final maximum number of copies includes:

[0082] Traversing and detecting the scaling flag value corresponding to each of the final maximum number of replicas in the recommendation array;

[0083] When it is detected that the scaling flag values ​​corresponding to the respective final maximum replica numbers in the recommended array are all greater than a preset value, or when there are values ​​that are both greater than and equal to the preset value, the smallest final maximum replica number is selected from the recommended array as the final target maximum replica number, and the final target maximum replica number is used to update the HPA maximum replica number in the HPA object;

[0084] When it is detected that the scaling flag values ​​corresponding to the respective final maximum replica numbers in the recommended array are all less than the preset value, or when there are values ​​that are both less than and equal to the preset value, the largest final maximum replica number is selected from the recommended array as the final target maximum replica number, and the final target maximum replica number is used to update the HPA maximum replica number in the HPA object.

[0085] Optionally, as in the aforementioned method, updating the maximum number of HPA replicas in the HPA object using the final target maximum number of replicas includes:

[0086] Detecting whether the final target maximum number of replicas is the same as the HPA maximum number of replicas;

[0087] When it is detected that the final target maximum number of replicas is the same as the HPA maximum number of replicas, not updating the HPA maximum number of replicas in the HPA object;

[0088] If it is detected that the final target maximum number of replicas is different from the HPA maximum number of replicas, detecting whether the final target maximum number of replicas exceeds a replica number threshold; wherein the replica number threshold is a preset multiple upper limit of the latest Pod number;

[0089] If it is detected that the final target maximum number of replicas does not exceed the replica number threshold, detecting whether the final target maximum number of replicas is less than the latest Pod number;

[0090] When it is detected that the final target maximum number of replicas is not less than the latest number of Pods, the maximum number of HPA replicas in the HPA object is replaced and updated based on the final target maximum number of replicas;

[0091] When it is detected that the final target maximum number of replicas is less than the latest number of Pods, a label value for enabling the scale-down label is obtained;

[0092] Detecting whether the tag value is a preset value;

[0093] When it is detected that the tag value is the preset value, replacing and updating the maximum number of HPA copies in the HPA object based on the final target maximum number of copies;

[0094] When it is detected that the label value is not the preset value, the maximum number of HPA replicas in the HPA object is replaced and updated based on the latest number of Pods;

[0095] When it is detected that the final target maximum number of replicas exceeds the replica number threshold, the HPA maximum number of replicas in the HPA object is replaced and updated based on the replica number threshold.

[0096] Optionally, as in the aforementioned method, the method further includes:

[0097] Querying a target scaling policy attribute value corresponding to the first indicator data from a scaling policy attribute value query table; wherein the scaling policy attribute value query table pre-stores a correspondence between an indicator data range of the first weight indicator and a scaling policy attribute value, wherein the scaling policy attribute value is used to characterize a scaling speed;

[0098] The scaling policy attribute value in the HPA object is updated using the target scaling policy attribute value.

[0099] In a second aspect, an embodiment of the present application provides a device for dynamic resource scaling and quota restriction of a cloud load balancer, the device comprising:

[0100] An acquisition module, configured to periodically poll and obtain business data and resource usage used to calculate the current maximum number of Pod replicas, wherein the business data includes first indicator data of a first weight indicator and second indicator data of a second weight indicator; wherein the first weight indicator and the second weight indicator are both business indicators that do not require adaptation to the HPA controller;

[0101] A first determining module, configured to determine a basic maximum number of replicas based on the first indicator data and the resource usage rate;

[0102] A second determination module is configured to dynamically adjust the basic maximum number of replicas through a hierarchical threshold trigger mechanism based on the second indicator data to determine a final maximum number of replicas;

[0103] An updating module is configured to update the maximum number of HPA copies in the HPA object using the final maximum number of copies.

[0104] In a third aspect, an embodiment of the present application provides an electronic device, which includes: a processor and a memory, the processor being used to execute a program for dynamic resource scaling and quota limitation of a cloud load balancer stored in the memory to implement the above-mentioned method for dynamic resource scaling and quota limitation of a cloud load balancer.

[0105] In a fourth aspect, an embodiment of the present application provides a storage medium, wherein the storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the above-mentioned method of dynamic resource scaling and quota limitation of the cloud load balancer.

[0106] The embodiments of the present application provide a method, device, electronic device, and storage medium for dynamic resource scaling and quota limiting of a cloud load balancer. The method includes: periodically polling to obtain business data and resource utilization used to calculate the current maximum number of Pod replicas, where the business data includes first indicator data of a first weight indicator and second indicator data of a second weight indicator; determining a basic maximum number of replicas based on the first indicator data and resource utilization; dynamically adjusting the basic maximum number of replicas through a hierarchical threshold trigger mechanism based on the second indicator data to determine the final maximum number of replicas; and updating the HPA maximum number of replicas in the HPA object using the final maximum number of replicas. This solution directly participates in the HPA replica number decision through the indicator data of business indicators, completely avoiding the architectural limitations of traditional solutions that must rely on adapters. Its core technological breakthrough lies in: calculating the final maximum number of replicas based on the indicator data of business indicators such as the first weight indicator and the second weight indicator, and being able to dynamically update the HPA maximum number of replicas in the HPA object. The entire process completely skips the intermediate adaptation link of converting business indicators into HPA native indicators, achieving seamless integration of business indicators and elastic scaling systems. This direct-connect design not only eliminates adapter development and maintenance costs, but also effectively avoids architectural redundancy and significantly reduces deployment configuration and version compatibility maintenance costs.

[0107] Furthermore, a hierarchical decision-making mechanism optimizes the calculation of the maximum number of replicas. This mechanism first generates a baseline maximum number of replicas based on resource utilization and a first weighted indicator, and then dynamically adjusts this value using a second weighted indicator. This multi-indicator collaborative scheduling design effectively overcomes the frequent pod number fluctuations caused by traditional independent triggering of each indicator, significantly improving scaling stability and further optimizing business performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0108] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.

[0109] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0110] Figure 1 A flowchart of an embodiment of a method for dynamic resource scaling and quota limitation of a cloud load balancer provided in an embodiment of the present application;

[0111] Figure 2 A flowchart of another embodiment of a method for dynamic resource scaling and quota restriction of a cloud load balancer provided in an embodiment of the present application;

[0112] Figure 3 A flowchart of another embodiment of a method for dynamic resource scaling and quota restriction of a cloud load balancer provided in an embodiment of the present application;

[0113] Figure 4 A flowchart of another embodiment of a method for dynamic resource scaling and quota restriction of a cloud load balancer provided in an embodiment of the present application;

[0114] Figure 5 A flowchart of another embodiment of a method for dynamic resource scaling and quota restriction of a cloud load balancer provided in an embodiment of the present application;

[0115] Figure 6 A flowchart of another embodiment of a method for dynamic resource scaling and quota restriction of a cloud load balancer provided in an embodiment of the present application;

[0116] Figure 7 A flowchart of another embodiment of a method for dynamic resource scaling and quota restriction of a cloud load balancer provided in an embodiment of the present application;

[0117] Figure 8 A block diagram of an embodiment of a device for dynamic resource scaling and quota limitation of a cloud load balancer provided in an embodiment of the present application;

[0118] Figure 9 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0119] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0120] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0121] References to "one embodiment" or "some embodiments" in this specification mean that a particular feature, structure, or characteristic described in conjunction with that embodiment is included in one or more embodiments of the present application. Thus, phrases such as "in one embodiment," "in some embodiments," "in other embodiments," and "in other embodiments" appearing in various places in this specification do not necessarily refer to the same embodiment, but rather mean "one or more but not all embodiments," unless otherwise specifically emphasized. The terms "including," "comprising," "having," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized.

[0122] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence refers to the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to achieve optimal results.

[0123] Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.

[0124] It should be understood that the size of the serial numbers of the steps in the following embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0125] To facilitate understanding of the embodiments of the present application, further explanation will be given below with reference to specific embodiments in conjunction with the accompanying drawings. The embodiments do not constitute a limitation on the embodiments of the present application.

[0126] The present invention provides a method for dynamic resource scaling and quota restriction of a cloud load balancer, which is applied to the HPAPLUS controller. Figure 1 , Figure 1 A flowchart of an embodiment of a method for dynamic resource scaling and quota limitation of a cloud load balancer provided in an embodiment of the present application. Figure 1 The process shown may include the following steps:

[0127] Step 101: Periodically poll to obtain business data and resource usage used to calculate the maximum number of Pod replicas. The business data includes first indicator data of a first weighted indicator and second indicator data of a second weighted indicator.

[0128] The first weighted indicator can directly reflect the service traffic pressure and is the primary basis for triggering expansion and contraction. Typical candidate indicators include: QPS or CPS (Connections Per Second, the number of new connections per second). Among them, QPS can reflect the request throughput actually processed by the system, is linearly correlated with user access behavior, and can directly reflect business load fluctuations; CPS characterizes the server's ability to establish connections and is particularly important for processing short-connection services. In this embodiment, QPS or CPS needs to be selected as the first weighted indicator based on the product characteristics (short-connection request-intensive or long-connection maintenance type).

[0129] Resource utilization is a comprehensive indicator that evaluates the node resource usage of Pods. It mainly includes the following three dimensions: CPU utilization, which reflects the real-time consumption of computing resources by Pods; memory utilization, which represents the memory usage level of Pods during runtime; and the number of Pods, which reflects the scale of Pod instances deployed on the node. These three dimensions together constitute the quantitative basis for resource pressure.

[0130] The second weight indicator serves as an auxiliary evaluation basis for service quality. Typical candidate indicators include: ACC or BPS (Bits PerSecond). Among them, ACC reflects the continuous connection pressure, but there is a lag (such as connection pool reuse causing indicator distortion); although BPS reflects network bandwidth occupancy, it fluctuates greatly due to the influence of data compression. Similarly, ACC or BPS needs to be selected as the second weight indicator based on product characteristics. In this embodiment, both the first weight indicator and the second weight indicator are business indicators that do not need to be adapted to the HPA controller, that is, in this application, there is no need to use an adapter to adapt the business indicators to the HPA controller, and the maximum number of copies of the Pod can be dynamically determined through collaborative analysis of the three types of data, so as to update and correct the maximum number of copies in the HPA object.

[0131] Step 102: Determine the basic maximum number of replicas based on the first indicator data and resource usage;

[0132] As can be seen from the above description, resource utilization, as a core quantitative indicator of node load, directly reflects the real-time stress status of the instance and is the key basis for triggering replica number adjustments to achieve load balancing. Furthermore, the first-weighted indicator, as the highest priority metric for business health, prioritizes scaling operations when it experiences anomalies. Based on this, by comprehensively analyzing the first-weighted indicator's first indicator data and resource utilization, we can more accurately calculate the base maximum number of replicas, thereby improving the accuracy of scaling decisions.

[0133] The above-mentioned maximum number of replicas represents the theoretical upper limit for capacity expansion, calculated by combining the first metric (e.g., QPS / CPS) with resource utilization. Simply put, this maximum number of replicas is calculated in real time by the HPA PLUS controller based on the current service load (QPS / CPS) and node resource headroom (CPU utilization, memory utilization, and number of Pods). It is an intermediate result that is not corrected by the second weighted metric (ACC / BPS).

[0134] Step 103: Dynamically adjust the basic maximum number of replicas based on the second indicator data through a hierarchical threshold trigger mechanism to determine the final maximum number of replicas;

[0135] Since the basic maximum number of replicas is calculated based only on traffic and resources, there may be deviations such as not taking into account the connection pool bottleneck (ACC) or ignoring the network bandwidth limit (BPS). Therefore, it is necessary to dynamically adjust the basic maximum number of replicas through the hierarchical threshold mechanism of the second weight indicator to form a more accurate expansion decision and obtain the final maximum number of replicas.

[0136] As can be seen from the above description, since the second-weighted indicators of ACC and BPS focus on the connection status and traffic dimensions respectively, and both have certain limitations, such indicator data can usually help correct the deviation in the maximum replica number decision.

[0137] Specifically, a multi-level threshold mechanism is set for the second weighted indicator. When different thresholds are triggered, the system will adopt the corresponding basic maximum replica number adjustment strategy. It should be noted that the default behavior of this strategy is to maintain the basic maximum replica number unchanged when no thresholds are triggered. In other words, the basic maximum replica number is determined as the final maximum replica number. The specific implementation details of this strategy will be elaborated in the subsequent examples and will not be explained in detail here.

[0138] Step 104: Update the HPA maximum number of copies in the HPA object using the final maximum number of copies.

[0139] The maximum number of HPA replicas is the maximum number of replicas determined by the HPA controller through elastic scaling based on the CPU / memory thresholds.

[0140] In this embodiment, the process of updating the HPA maximum number of replicas in the HPA object using the final maximum number of replicas is: the HPAPLUS controller directly modifies the maxRepl icas field of the HPA object through the Kubernetes (k8s) API (Application Programming Interface) to use the final maximum number of replicas to correct the HPA maximum number of replicas determined by the HPA controller, and when it is confirmed that the new configuration has taken effect and the scaling behavior is as expected, the HPA PLUS controller will immediately recalculate the final maximum number of replicas according to the method provided in this embodiment.

[0141] This solution directly utilizes business metric data to determine the number of HPA replicas, completely avoiding the architectural limitations of traditional solutions that rely on adapters. Its core technological breakthrough lies in calculating the final maximum number of replicas based on the metric data of business metrics, such as the first and second weighted metrics, and dynamically updating the maximum number of HPA replicas in the HPA object. This entire process completely bypasses the intermediate adaptation step of converting business metrics into native HPA metrics, achieving seamless integration between business metrics and the elastic scaling system. This direct-connect design not only eliminates adapter development and maintenance costs, but also effectively avoids architectural redundancy, significantly reducing deployment configuration and version compatibility maintenance costs.

[0142] Furthermore, this technical solution optimizes the calculation of the maximum number of replicas through a hierarchical decision-making mechanism. It first generates a baseline maximum number of replicas based on resource utilization and a first weighted indicator, and then dynamically adjusts this value using a second weighted indicator. This multi-indicator collaborative scheduling design effectively overcomes the frequent pod number fluctuations caused by traditional independent triggering of each indicator, significantly improving scaling stability and further optimizing business performance.

[0143] like Figure 2 As shown, as an optional implementation, as in the aforementioned method, the step 101 of periodically polling to obtain the business data and resource usage used to calculate the current maximum number of Pod replicas includes the following steps:

[0144] Step 201, obtaining a first resource usage rate sequence currently collected using a time window;

[0145] A sliding time window mechanism is required (the default time window size is 5 sampling points, which is configurable) to collect the latest data:

[0146] 1. Record CPU utilization and memory utilization at fixed intervals to form a time series. 2. Synchronously obtain the number of Pods in the Ready state in the Kubernetes cluster at each sampling time.

[0147] The above CPU utilization, memory utilization and the number of Pods at the corresponding time points are aligned in time and space to construct the first resource utilization sequence (ResourceUtilizationRatio1[Ti....Tj]) containing multi-dimensional indicator data for the HPAPLUS controller to obtain.

[0148] Step 202, extract the start and end timestamps of the time window;

[0149] In order to query and obtain the first indicator data and the second indicator data of the sequence, in this embodiment, it is necessary to obtain the start timestamp and the end timestamp accurate to the second level.

[0150] Step 203, determining whether the time span of the start and end timestamps is less than a preset time span;

[0151] The time span is determined by the time difference between the start timestamp (startTime) and the end timestamp (endTime). The preset time span can be set according to actual needs and is not limited here, for example, it can be set to 300 seconds, 350 seconds, 400 seconds, etc.

[0152] When it is determined that the time span of the start and end timestamps is less than the preset time span, in order to avoid monitoring value lag, resulting in too little first indicator data and second indicator data being queried, step 204 can be executed to obtain the first data sequence; when it is determined that the time span of the start and end timestamps is equal to or greater than the preset time span, step 205 can be executed to obtain the first data sequence.

[0153] Step 204 , querying the cluster monitoring system based on a preset time span to obtain a current first data sequence;

[0154] If endTime-startTime<300 (preset time span), the time query range is directly extended to 300 seconds, that is, startTime=endTime-300. Based on the extended time span, that is, the preset time span, multiple first indicator data and multiple second indicator data are queried from the cluster monitoring system (such as Prometheus) to obtain. Among the multiple first indicator data and second indicator data obtained, the monitoring points with a threshold value less than the initial threshold (for example, the default value of QPS: 1000) are filtered out. This threshold value is determined by the initial parameters and is generally kept at the default value. At the same time, if too much first indicator data and second indicator data are queried, the first indicator data and second indicator data of up to the last 8 (default value, adjustable) time points can be selected to construct the current first data sequence (Value1[Tn....Tm]) for acquisition by the HPA PLUS controller.

[0155] Step 205 , querying the cluster monitoring system based on the time span of the start and end timestamps to obtain the current first data sequence;

[0156] After obtaining multiple first indicator data and second indicator data based on the time span query of the start and end timestamps, threshold screening and excessive data removal processing are also required to construct the current first data sequence for acquisition by the HPAPLUS controller.

[0157] Step 206: Time-align the first resource usage sequence and the first data sequence to determine a second resource usage sequence and a second data sequence;

[0158] At present, two sequences ResourceUtilizationRatio1[Ti....Tj] and Value1[Tn....Tm] are obtained. The time points of the two sequences are intersected to obtain the common time range [Tx...Ty]. The resource utilization sequence in this range is taken as the second resource utilization sequence ResourceUtilizationRatio2 for subsequent calculations. Similarly, multiple first indicator data and multiple second indicator data close to this range are selected from Value1 as the second data sequence Value2 for subsequent calculations.

[0159] Step 207 : Obtain resource usage based on the second resource usage sequence, and obtain first indicator data of the first weight indicator and second indicator data of the second weight indicator based on the second data sequence.

[0160] As can be seen from the above description, resource utilization includes the current number of Pods, current memory utilization, and current CPU utilization. The process of obtaining resource utilization based on the second resource utilization sequence can be implemented through steps A1 to A7:

[0161] Step A1: Obtain the latest number of Pods from the second resource usage sequence;

[0162] Specifically, you can use the following call statement to obtain the latest number of Pods from the second resource usage sequence:

[0163] lastPodNum = resourceUtilizationRatio2[len(resourceUtilizationRatio2)-1].countPods. len(resourceUtilizationRatio2)-1 indicates the index of the last element in the second resource utilization sequence, resourceUtilizationRatio2[...] indicates accessing the last element in the second resource utilization sequence by index, and .countPods indicates accessing the countPods property or field in the second resource utilization sequence to obtain the latest number of Pods.

[0164] Step A2: determining the average number of Pods and the standard deviation of Pods based on the second resource usage sequence;

[0165] The average number of Pods can reflect the long-term load level, and the Pod standard deviation can quantify the volatility of the number of Pods.

[0166] Specifically, when determining the average number of Pods, it is necessary to obtain the total number of Pods totalPods and the resource statistics dimension resourceLength (for example, the length of the time period) from the second resource utilization sequence, and use the formula averagePods = totalPods / resourceLength to calculate the average number of Pods in each unit dimension (such as per hour).

[0167] When determining the Pod standard deviation, it can be achieved through the following statement: stdevPods=STDEV(resourceUtilizationRatio2.countPods).

[0168] Among them, resourceUtilizationRatio2.countPods means extracting the countPods field values ​​of all elements from the second resource utilization sequence to form a Pod quantity sequence. Example: If the array contains 3 records, and their countPods values ​​are [10, 15, 12], then the extracted sequence is [10, 15, 12]. STDEV() represents the standard deviation function (StandardDeviation), which is used to calculate the discrete degree of the above extracted sequence to obtain the Pod standard deviation.

[0169] Step A3: Determine whether the 2σ standard deviation noise filtering mechanism is met based on the latest Pod number, average Pod number, and Pod standard deviation;

[0170] Use a 2σ standard deviation noise filtering mechanism to detect whether the latest Pod number is a noise point. If the 2σ standard deviation noise filtering mechanism is met, that is, |lastPodNum - averagePods| ≤ 2*stdevPods, the latest Pod number is not a noise point, and step A4 is executed. If the 2σ standard deviation noise filtering mechanism is not met, that is, |lastPodNum - averagePods| > 2*stdevPods, the latest Pod number is a noise point, and step A5 is executed.

[0171] Step A4: Determine the latest number of Pods as the current number of Pods;

[0172] Step A5: The latest number of Pods at the previous running time is determined as the current number of Pods;

[0173] When determining that the latest Pod number is a noise point, it is necessary to roll back to the Pod number at the previous runtime [currentPods = resourceUtilizationRatio2[len(resourceUtilizationRatio2)-2].countPods)]. The above process is a specific process for smoothing the Pod number. Its core purpose is to eliminate the interference caused by Pod number fluctuations through technical means to prevent outliers from affecting the stability and reliability of scaling decisions.

[0174] Step A6: obtaining the latest memory utilization from the second resource utilization sequence, and determining the latest memory utilization as the current memory utilization;

[0175] Similarly, the latest memory utilization can be obtained from the second resource usage sequence by calling the following statement:

[0176] latestMemory=resourceUtilizationRatio2[len(resourceUtilizationRatio2)-1].MemoryUtilizationRatio.

[0177] Wherein, len(resourceUtilizationRatio2)-1 indicates obtaining the index of the last element of the second resource utilization sequence, resourceUtilizationRatio2[...] indicates accessing the last element of the second resource utilization sequence by index, and .MemoryUtilizationRatio indicates accessing the memory attribute or field of the second resource utilization sequence to obtain the latest memory utilization.

[0178] After obtaining the latest memory utilization, perform numerical conversion on it to obtain the current memory utilization. The numerical conversion process can be achieved through the following statement:

[0179] currentMemory=float64(latestMemory) / 100.0.

[0180] Step A7: Obtain the latest CPU utilization from the CPU utilization sequence, and determine the latest CPU utilization as the current CPU utilization.

[0181] Similarly, the latest memory utilization can be obtained from the second resource usage sequence by calling the following statement:

[0182] latestCpu=resourceUtilizationRatio2[len(resourceUtilizationRatio2)-1].CpuUtilizationRatio.

[0183] In this example, len(resourceUtilizationRatio2)-1 indicates obtaining the index of the last element of the second resource utilization sequence, resourceUtilizationRatio2[...] indicates accessing the last element of the second resource utilization sequence by index, and .CpuUtilizationRatio indicates accessing the CPU attribute or field of the second resource utilization sequence to obtain the latest CPU utilization.

[0184] After obtaining the latest CPU utilization, perform numerical conversion on it to obtain the current CPU utilization. The numerical conversion process can be achieved through the following statement:

[0185] currentCpu=float64(latestCpu) / 100.0.

[0186] The above process of obtaining the first indicator data of the first weight indicator and the second indicator data of the second weight indicator based on the second data sequence can be implemented through steps B1 to B2:

[0187] Step B1, obtaining the latest first indicator data from the second data sequence, and determining the latest first indicator data as the first indicator data of the first weighted indicator;

[0188] Specifically, the following call statement can be used to obtain the latest first indicator data from the second data sequence (taking QPS as the first weighted indicator as an example):

[0189] currentQps=Value 2[countQps-1].Value; wherein, Value 2 represents the second data sequence, in which multiple QPS data recorded in chronological order are stored. Each element contains a Value attribute (indicating the QPS value), countQps-1 is the index of the latest data (array subscripts start at 0), and .Value represents extracting the Value attribute of the element at the index position, the specific QPS value, and the latest QPS obtained is determined as the first indicator data.

[0190] Step B2: Obtain the latest second indicator data from the second data sequence, and determine the latest second indicator data as the second indicator data of the second weighted indicator.

[0191] Similarly, the latest second indicator data can be obtained from the second data series through the following call statement (taking ACC as the second weight indicator as an example):

[0192] currentAcc = Value 2[countAcc-1].Value; where Value 2 represents the second data sequence, which also stores multiple ACC data items recorded in chronological order. Each element contains a Value attribute (representing the ACC value). CountQps-1 is the index of the most recent data item (array subscripts start at 0). .Value represents extracting the Value attribute of the element at that index, specifically the ACC value. The most recent ACC value obtained is determined as the second indicator data. It should be noted that the acquisition logic for the first and second indicator data is independent of each other and can be acquired in parallel or sequentially, which is not limited here.

[0193] In all the following embodiments, QPS is used as the first weight indicator to illustrate the complete calculation process, and the calculation process of CPS as the first weight indicator is the same, and the calculation process of CPS is not described in detail here.

[0194] like Figure 3 As shown, as an optional implementation, as in the aforementioned method, step 102 of determining the basic maximum number of replicas based on the first indicator data and resource usage includes the following steps:

[0195] Step 301: Obtain a maximum indicator threshold, a memory utilization threshold, and a CPU utilization threshold of a first weighted indicator;

[0196] The maximum metric threshold indicates the maximum QPS value set for the instance specification. The CPU utilization threshold (a.cpuRatio) and memory utilization threshold (a.memoryRatio) are set by the corresponding instance specification parameters, generally 0.9 and 0.4, respectively. They represent the maximum CPU and memory utilization thresholds allowed for a single Pod to operate in the long term.

[0197] Step 302: Determine the maximum number of CPU replicas based on the current number of Pods, the maximum indicator threshold, the first indicator data, the CPU utilization threshold, and the current CPU utilization;

[0198] The maximum number of CPU replicas can be calculated using the following formula:

[0199]

[0200] cpuMaxReplicas indicates the maximum number of CPU replicas, currentPods indicates the current number of Pods, targetQps indicates the maximum indicator threshold, currentQps indicates the current QPS (i.e., the first indicator data), a.cpuRatio indicates the CPU utilization threshold, and currentCpu indicates the current CPU utilization.

[0201] Step 303: Determine the maximum number of memory replicas based on the current number of Pods, the maximum indicator threshold, the first indicator data, the memory utilization threshold, and the current memory utilization;

[0202] The maximum number of memory copies can be calculated using the following formula:

[0203]

[0204] Where memoryMaxReplicas indicates the maximum number of memory replicas, currentPods indicates the current number of Pods, targetQps indicates the maximum indicator threshold, currentQps indicates the current QPS (i.e., the first indicator data), a.memoryRatio indicates the memory utilization threshold, and currentMemory indicates the current memory utilization.

[0205] Step 304 : Select the maximum value from the maximum number of CPU replicas and the maximum number of memory replicas, and round up the maximum value to obtain the basic maximum number of replicas.

[0206] maxReplicas = max(cpuMaxReplicas, memoryMaxReplicas), which takes the maximum of the calculated values ​​of the maximum number of CPU replicas and the maximum number of memory replicas. Then, maxReplicasInt32 = int32(math.Ceil(maxReplicas)), after numerical conversion and rounding up, obtains the basic maximum number of replicas.

[0207] like Figure 4 As shown, as an optional implementation, as in the aforementioned method, step 103 dynamically adjusts the basic maximum number of copies based on the second indicator data through a hierarchical threshold trigger mechanism to determine the final maximum number of copies, including the following steps:

[0208] Step 401: Determine a per-core unit value based on the second indicator data and the current number of Pods;

[0209] Determine the total number of cores based on the current number of Pods, and then divide the second indicator data threshold by the total number of cores to obtain the per-core unit value.

[0210] Step 402 , detecting whether the per-core unit value exceeds a preset indicator data threshold range;

[0211] The preset indicator data threshold range consists of a pre-set pressure value and limit value for each core. The pressure value and limit value are pre-set based on the type and specifications of different listeners. The purpose of setting the preset indicator data threshold range is to enable hierarchical threshold triggering to dynamically adjust the basic maximum number of replicas.

[0212] Specifically, when it is detected that the per-core unit value is within the preset indicator data threshold range, it is necessary to cooperate with the first weight indicator and decide whether to increase the number of Pods according to different strategies, and steps 403 to 405 need to be executed; when it is detected that the per-core unit value exceeds the preset indicator data threshold range, the maximum number of replicas is immediately modified, that is, the number of Pods is increased, and step 406 needs to be executed; when it is detected that the per-core unit value is lower than the preset indicator data threshold range, the impact of the second weight indicator on the basic maximum number of replicas is ignored, and step 407 needs to be executed.

[0213] Step 403: Obtain the real-time change rate of the first weight index;

[0214] The real-time rate of change can directly reflect the dynamic fluctuation characteristics of the first weighted indicator. In particular, when the second indicator data is above the pressure value but below the limit value, the rate of change can accurately quantify the abnormal fluctuation trend of the first weighted indicator. This correlation analysis provides data support for whether to increase the number of pods.

[0215] Step 404, detecting whether the real-time change rate exceeds a preset change rate threshold;

[0216] The preset change rate threshold is a pre-set response boundary value used to determine whether the fluctuation of the first weight indicator triggers the intervention mechanism; the specific preset change rate threshold can be set according to actual needs and is not limited here.

[0217] When it is detected that the real-time change rate exceeds the preset change rate threshold, it indicates that the current traffic has exceeded the carrying capacity of the basic maximum number of replicas, and step 405 needs to be executed immediately to adjust the number.

[0218] When it is detected that the real-time change rate does not exceed the preset change rate threshold, it indicates that the current traffic does not exceed the carrying capacity of the basic maximum number of replicas, and step 407 needs to be executed immediately without adjusting the number.

[0219] Step 405: Adjust the basic maximum number of replicas according to the preset number of adjusted Pods, and determine the adjusted basic maximum number of replicas as the final maximum number of replicas.

[0220] Specifically, you can increase the preset number of adjusted Pods based on the basic maximum number of replicas, and determine the basic maximum number of replicas for the increased Pods as the final maximum number of replicas.

[0221] For example, if the basic maximum number of replicas is 10 and the preset number of adjusted Pods is 2 (which can be adjusted), the final maximum number of replicas is 12.

[0222] Step 406: Modify the basic maximum number of replicas based on the second indicator data, and determine the modified basic maximum number of replicas as the final maximum number of replicas.

[0223] The specific data modification process can be achieved through steps C1 to C3:

[0224] Step C1, determining the required number of cores based on the per-core unit value and the limit value;

[0225] Divide the per-core unit value by the limit value and round up the result to determine the required number of cores. For example, if the per-core unit value is 500 and the limit value is 200, the required number of cores is 3.

[0226] Step C2: Determine the number of replicas to be added based on the required number of cores;

[0227] In this embodiment, each Pod is assumed to have the same number of cores. Therefore, the number of replicas added is determined by the required number of cores and the number of cores in the Pod. For example, if a Pod is assumed to have 2 cores, the number of replicas added is 2 because the required number of cores is 3. If a Pod is assumed to have 3 or more cores, the number of replicas added is 1 because the required number of cores is 3.

[0228] Step C3: Modify the basic maximum number of copies based on the number of copies added.

[0229] Add the number of replicas added to the basic maximum number of replicas, and the result is the final maximum number of replicas.

[0230] Step 407: determine the basic maximum number of copies as the final maximum number of copies.

[0231] When the per-core unit value is lower than the stress value, or the real-time change rate does not exceed the preset change rate threshold, it indicates that the current traffic does not exceed the carrying capacity of the basic maximum number of replicas. There is no need to adjust the number of baseline maximum replicas or modify the data. Determining the basic maximum number of replicas as the final maximum number of replicas can meet business performance requirements.

[0232] This method of optimizing the calculation of the maximum number of replicas through a multi-indicator collaborative decision-making mechanism effectively overcomes the frequent fluctuations in the number of Pods caused by the independent triggering of traditional indicators, significantly improves the stability of scaling, and further optimizes business performance.

[0233] In actual applications, after determining the final maximum number of replicas, the HPA object attributes will not be updated directly at this time. Instead, it will be placed in the recommendation array, which is the election window. The recommendation array usually retains historical recommendation values, and all the recommendation values ​​in the recommendation array are used to decide whether to update the HPA object attributes. This can effectively avoid frequent fluctuations in the number of replicas caused by instantaneous indicator fluctuations. Therefore, before using the final maximum number of replicas to update the HPA maximum number of replicas in the HPA object, it is necessary to execute Figure 5 The embodiment shown, as Figure 5 The specific steps shown are as follows:

[0234] Step 501: Determine the scaling flag value corresponding to the final maximum number of replicas based on the first indicator data, the current number of Pods, the current CPU utilization, the current memory utilization, the maximum indicator threshold, the CPU utilization threshold, the memory utilization threshold, and the maximum number of HPA replicas.

[0235] The scaling flag value (isDown) is a key status flag used to identify the direction of scaling decisions. Specifically, a scaling flag value of 0 indicates scaling up; a scaling flag value greater than 1 indicates scaling down; and a scaling flag value of 1 indicates remaining unchanged.

[0236] The specific process of determining the value of the telescopic flag can be achieved through steps D1 to D4:

[0237] Step D1: determining the CPU utilization of a single Pod based on the first indicator data, the current number of Pods, and the current CPU utilization; and determining the memory utilization of a single Pod based on the first indicator data, the current number of Pods, and the current memory utilization;

[0238] Calculate the maximum QPS that each Pod can currently handle based on CPU / memory, that is, single-Pod CPU utilization / single-Pod memory utilization.

[0239] The CPU utilization of a single Pod can be calculated using the following formula:

[0240] perPodCpuRate=currentQps / (currentPods×currentCpu);

[0241] Among them, peerPodCpuRateb indicates the CPU utilization of a single Pod.

[0242] The memory utilization of a single Pod can be calculated using the following formula:

[0243] perPodMemoryRate=currentQps / (currentPods×currentMemory);

[0244] Among them, perPodMemoryRate indicates the memory utilization of a single Pod.

[0245] Step D2: determining the memory Pod incremental requirement value based on the single Pod memory utilization, the first indicator data, and the maximum indicator threshold, and determining the CPU Pod incremental requirement value based on the single Pod CPU utilization, the first indicator data, and the maximum indicator threshold;

[0246] The memory Pod incremental demand value / CPU Pod incremental demand value is used to indicate how many Pods are needed for the memory or CPU to reach the currently set QPS.

[0247] The incremental memory Pod requirement can be calculated using the following formula:

[0248] memoryPods=(currentQps-targetQps) / perPodMemoryRate;

[0249] Among them, memoryPods indicates the incremental demand value of memory Pod.

[0250] The CPUPod incremental requirement value can be calculated using the following formula:

[0251] cpuPods=(currentQps-targetQps) / perPodCpuRate;

[0252] cpuPods indicates the incremental demand value of CPUPod.

[0253] Step D3: Determine the remaining available Pod capacity of the CPU based on the CPU utilization threshold, the current CPU utilization, the current number of Pods, and the maximum number of HPA replicas; and determine the remaining available Pod capacity of the memory based on the memory utilization threshold, the current memory utilization, the current number of Pods, and the maximum number of HPA replicas.

[0254] The CPU remaining available Pod capacity value / memory remaining available Pod capacity value is used to indicate the number of CPU / memory Pods available based on the maximum number of HPA replicas.

[0255] The remaining available Pod capacity of the CPU can be calculated using the following formula:

[0256] remainderCpuPods=(a.cpuratio-currentCpu)×currentPods+

[0257] oldMaxReplicas–currentPods;

[0258] RemainderCpuPods indicates the remaining available Pod capacity of the CPU, and oldMaxRepl icas indicates the maximum number of HPA replicas.

[0259] The remaining available Pod capacity in memory can be calculated using the following formula:

[0260] remainderMemoryPods=(a.memoryratio-currentMemory)×currentPods+

[0261] oldMaxReplicas-currentPods;

[0262] RemainderMemoryPods indicates the remaining available Pod capacity in memory.

[0263] Step D4: Determine the scaling flag value corresponding to the final maximum number of replicas based on the final maximum number of replicas, the maximum number of HPA replicas, the memory Pod incremental requirement value, the CPU Pod incremental requirement value, the CPU remaining available Pod capacity value, the memory remaining available Pod capacity value, the first indicator data, the current CPU utilization rate, the current number of Pods, and the latest number of Pods.

[0264] The process of determining the scaling flag value corresponding to the final maximum number of replicas based on the above data may be implemented through steps E1 to E18:

[0265] Step E1: Check whether the final maximum number of replicas is the same as the latest number of Pods;

[0266] Step E2: If it is detected that the final maximum number of replicas is the same as the latest number of Pods, check whether the final maximum number of replicas is selected from the maximum number of CPU replicas;

[0267] If the final maximum number of replicas is the same as the latest number of Pods, this means that the number of Pods did not change within the [Tx...Ty] time window. This situation is divided into two categories based on whether the basic maximum number of replicas is based on CPU or memory. Each category is further divided into two cases. There are a total of four cases, namely the first to fourth conditions below. In all four cases, the isDown scaling flag value is set to 3, and the final maximum number of replicas remains the latest number of Pods.

[0268] Step E3: When it is detected that the final maximum number of replicas is selected from the maximum number of CPU replicas, determine whether the first condition or the second condition is met based on the CPUPod increment requirement value and the CPU remaining available Pod capacity value;

[0269] The first condition is: cpuPods>=0&&cpuPods<=1.0&&remainderCpuPods<=1.0

[0270] If the first condition is met, it means that the current currentQps just exceeds the target Qps maximum threshold, the number of Pods that exceed the target is less than 1, and the number of remaining available Pods is also less than 1.

[0271] The second condition is: cpuPods<0&&remainderCpuPods>=math.Abs(cpuPods)&&remainderCpuPods<=math.Abs(cpuPods)+2; where math.Abs(cpuPods) represents the absolute value of the CPUPod incremental demand value.

[0272] If the second condition is met, it means that the current currentQps has not yet reached the target Qps threshold, and the remaining available Pods can cover the required number of Pods, with no more than two Pods exceeding the target number.

[0273] Step E4: if the first condition or the second condition is met, setting the value of the telescopic flag to the first value;

[0274] When the first condition or the second condition is met, it indicates that scaling is required. Therefore, the scaling flag value may be set to the first data (eg, 3), ie, isDown=3.

[0275] Step E5: When it is detected that the final maximum number of replicas is selected from the maximum number of memory replicas, determine whether the third condition or the fourth condition is met based on the memory Pod increment requirement value and the memory remaining available Pod capacity value;

[0276] The third condition is:

[0277] memoryPods>=0&&memoryPods<=1.0&&remainderMemoryPods<=1.0

[0278] The explanation is the same as the first condition above. The calculated value comes from the memory and will not be repeated here.

[0279] The fourth condition is:

[0280] memoryPods<0&&remainderMemoryPods>=math.Abs(memoryPods)&&remainderMemoryPods<=math.Abs(memoryPods)+2;

[0281] The explanation is the same as the second condition above. The calculated value comes from the memory and is not repeated here.

[0282] Step E6: If the third condition or the fourth condition is met, set the value of the telescopic flag to the first value;

[0283] When the third condition or the fourth condition is met, it indicates that scaling down is also required. Therefore, the scaling flag value can also be set to the first value, ie, isDown=3.

[0284] Step E7: When it is detected that the final maximum number of replicas is different from the latest number of Pods, determine the average data of the first indicator based on the second data sequence;

[0285] That is, each first indicator data is obtained from the second data sequence, and then all the first indicator data are averaged to obtain the first indicator average data.

[0286] Step E8: Determine whether the fifth condition is satisfied based on the first indicator data and the first indicator average data, or determine whether the sixth condition is satisfied based on the current number of Pods and the current number of Pods at the previous running time.

[0287] The fifth condition is: the first indicator data is less than the average data of the first indicator calculated last time, and the decline is greater than 5%;

[0288] The sixth condition is that the currentPods value must be less than 2 compared to the currentPods value at the previous runtime.

[0289] Step E9: If the fifth condition or the sixth condition is met, set the value of the telescopic flag to the second value;

[0290] When the fifth condition or the sixth condition is met, it indicates that scaling is also required. Therefore, the scaling flag value can also be set to the second data (such as 4), that is, isDown=4.

[0291] The fifth and sixth conditions also include a special case: when client request pressure is high enough and the CPU usage of all pods is almost fully utilized, but the maximum target QPS is still far from being reached, the monitoring results in oscillation around a certain value. Due to the logic setting of isDown = 4, capacity expansion cannot be carried out in a timely manner for a period of time. Therefore, the following operation is used to optimize and handle this situation.

[0292] During each calculation, the current CPU utilization in the last resource monitoring point ResourceUtilizationRatio2 is taken to determine whether it is greater than 80%. If it is, record.highPressure++ is recorded. If it is less than 80% only once, record.highPressure=0 is set. If record.highPrressure>3 appears, the logic of setting isDown=4 is skipped and the subsequent isDowns value setting logic is continued.

[0293] Step E10: If the fifth and sixth conditions are not met, detecting whether the current CPU utilization is less than a preset pressure threshold;

[0294] The preset pressure threshold (such as 0.55) can be set according to actual needs and is not set here.

[0295] Step E11: When it is detected that the current CPU utilization is less than the preset pressure threshold, the scaling flag value is set to a third value;

[0296] When it is detected that the current CPU utilization is less than the preset pressure threshold, it means that the current pressure is not enough, and the scaling flag value is set to a third value (such as 5), that is, isDown=5.

[0297] In practice, if the current CPU utilization is less than the preset pressure threshold, or if the fifth or sixth condition is met, the stopScaleOut flag is applied. If the calculated final maximum number of replicas is greater than or equal to the HPA maximum number of replicas, scaling is stopped. This is because when QPS drops significantly over a short period of time, the native HPA downscale requires a 5-minute wait (set by the --horizontal-pod-autoscaler-downscale-stabilization-window parameter) before it begins. The resulting currentPods remain largely unchanged, causing a significant decrease in the calculated perPodCpuRate and perPodMemoryRate, while the corresponding final maximum number of replicas increases significantly. If the previous QPS was just around the target QPS (e.g., 50k), a further surge in traffic within the downscale window could easily break through the 50k QPS limit, directly pushing the QPS to 60k-70k. After a period of stabilization, the calculated final maximum number of replicas will drop again, but will be less than the latest number of Pods at that time, preventing active downscaling and forcing it to remain at its original value. At this time, the monitoring value lossPods = the final maximum number of replicas - the latest number of Pods, indicating the number of Pods lost at this time.

[0298] In the above process, the calculated perPodMemoryRate value drops most significantly. This is because when QPS pressure decreases, CPU utilization may decrease, but the average memory utilization remains basically unchanged. Since scaling down requires a certain period of stabilization, the currentPods value remains basically unchanged, resulting in a very large calculated memoryMaxReplicas, which in turn increases the final maximum number of replicas. Marking it with stopScaleOut and stopping scaling are optimization methods.

[0299] Step E12: If it is detected that the current CPU utilization is not less than the preset pressure threshold, check whether the final maximum number of replicas is less than the latest number of Pods;

[0300] In step E13, if it is detected that the final maximum number of replicas is less than the latest number of Pods, the scaling flag value is set to the fourth value;

[0301] If the final maximum number of replicas is detected to be less than the latest number of Pods, it means that after the Pod expansion is completed, the QPS has stabilized and the achieved value exceeds the target QPS. The calculated perPodCpuRate and perPodMemoryRate will also tend to a larger value. As a result, the final maximum number of replicas calculated later will be smaller than the latest number of Pods. Since active scaling will not be performed, the scaling flag value is set to the fourth value (such as 6), that is, isDown = 6.

[0302] Step E14: When it is detected that the final maximum number of replicas is greater than the latest number of Pods, detect whether the first indicator data is less than a preset initial threshold value;

[0303] The preset initial threshold value is the initial threshold value (1000) mentioned in the above embodiment. This value can be set according to actual needs and is not limited here.

[0304] Step E15: When it is detected that the first indicator data is less than the preset initial threshold, the scaling flag value is set to a fifth value;

[0305] When it is detected that the first indicator data is less than the preset initial threshold, it indicates that scaling down is required, so the scaling flag value is set to the fifth value (eg, 2), ie, isDown=2.

[0306] Step E16: If it is detected that the first indicator data is not less than the preset initial threshold, detect whether the final maximum number of replicas is greater than the HPA maximum number of replicas;

[0307] In situations that do not fall under the above categories, you need to compare the final maximum number of replicas with the HPA maximum number of replicas to set the scaling flag data.

[0308] Step E17: If it is detected that the final maximum number of replicas is not greater than the maximum number of HPA replicas, the scaling flag value is set to a sixth value.

[0309] When the final maximum number of replicas is less than or equal to the maximum number of HPA replicas, the description remains unchanged, so the scaling flag value needs to be set to the sixth value (such as 1), that is, isDown=1.

[0310] Step E18: When it is detected that the final maximum number of replicas is greater than the maximum number of HPA replicas, the scaling flag value is set to the seventh value.

[0311] When the final maximum number of replicas is greater than the maximum number of HPA replicas, it indicates that capacity expansion is required, so the scaling flag value needs to be set to the seventh value (such as 0), that is, isDown=0.

[0312] Step 502: Store the final maximum number of replicas and the scaling flag value corresponding to the final maximum number of replicas into a recommendation array.

[0313] The recommended array also includes the final maximum number of replicas determined at multiple historical moments, as well as the corresponding scaling flag values ​​for each final maximum number of replicas. In practice, changes are only made when the isDown flag type of all recommended values ​​within the election window is identical (all <= 1 or all >= 1) to avoid misoperation caused by transient indicator fluctuations.

[0314] Based on the above description, if Figure 6 As shown, as an optional implementation, as in the aforementioned method, step 104 of updating the maximum number of HPA copies in the HPA object using the final maximum number of copies includes the following steps:

[0315] Step 601: traverse and detect the scaling flag value corresponding to each final maximum number of replicas in the recommendation array;

[0316] Traverse all the final maximum replica numbers in the recommended array and check the corresponding scaling flag values. This step is used to collect the scaling status of all replicas and provide a data basis for subsequent decision-making.

[0317] Step 602: If it is detected that the scaling flag values ​​corresponding to the final maximum replica numbers in the recommended array are all greater than the preset value, or if there are both values ​​greater than and equal to the preset value, the smallest final maximum replica number is selected from the recommended array as the final target maximum replica number, and the final target maximum replica number is used to update the HPA maximum replica number in the HPA object.

[0318] When all scaling flag values ​​are greater than the preset value (strictly greater than), or when some scaling flag values ​​are equal to the preset value and others are greater than the preset value (mixed state), the smallest final maximum number of replicas is selected from the recommended array and the HPA maximum number of replicas in the HPA object is updated. This selection is intended to prioritize the minimum number of replicas to control costs when resource utilization is high.

[0319] In step 603, if it is detected that the scaling flag values ​​corresponding to the final maximum replica numbers in the recommended array are all less than the preset value, or if there are both values ​​less than and equal to the preset value, the largest final maximum replica number is selected from the recommended array as the final target maximum replica number, and the final target maximum replica number is used to update the HPA maximum replica number in the HPA object.

[0320] When all scaling flag values ​​are less than the preset value (strictly less than), or when some scaling flag values ​​are equal to the preset value and others are less than the preset value (mixed state), the largest final maximum replica count is selected from the recommended array and the HPA maximum replica count in the HPA object is updated. This selection is intended to ensure service availability by selecting the maximum replica count when resource utilization is low.

[0321] In actual application, when the preset value is 1, the process for selecting the final maximum number of replicas is as follows: within the election window, if the scaling flag value isDown for all recommended values ​​(i.e., the final maximum number of replicas) meets one of the following two conditions, the corresponding selection operation is performed; otherwise, the selection update is not performed (all values ​​are equal to 1, or one value is greater than 1 and the other is less than 1).

[0322] 1. When isDown <= 1, select the maximum final maximum number of replicas of all values ​​for the HPA maximum number of replicas update;

[0323] 2. When isDown>=1, select the minimum final maximum number of replicas of all values ​​for the HPA maximum number of replicas update.

[0324] The above strategy for selecting the final maximum number of replicas is: select the maximum when expanding and the minimum when shrinking, to prevent the operation delay of the election window from affecting the overall response speed. The strategy can also be flexibly adjusted based on the characteristics of the business itself, and there is no limit on this.

[0325] The specific process of updating the maximum number of HPA copies in the HPA object using the final target maximum number of copies in steps 602 and 603 can be implemented through steps F1 to F10:

[0326] Step F1, checking whether the final target maximum number of replicas is the same as the HPA maximum number of replicas;

[0327] If it is detected that the final target maximum number of replicas is the same as the HPA maximum number of replicas, step F2 is immediately executed; if it is detected that the final target maximum number of replicas is different from the HPA maximum number of replicas, step F3 is immediately executed.

[0328] Step F2: The maximum number of HPA copies in the HPA object is not updated;

[0329] If the selected final target maximum number of replicas is equal to the HPA maximum number of replicas, no update is performed, reducing the number of modifications to the HPA object to save unnecessary interface calls and system overhead.

[0330] Step F3, detecting whether the final target maximum number of copies exceeds the copy number threshold;

[0331] The replica count threshold is a preset upper limit of the latest Pod count. This preset multiple can be set based on actual needs and is not specified here. For example, during stress testing and traffic flow, there may be periods of time when the calculated maximum memory replica count or maximum CPU replica count is very large. This is generally caused by the Pods having just been expanded and not yet evenly distributing traffic. Therefore, the final target maximum replica count for each adjustment is limited to no more than twice the latest Pod count.

[0332] If it is detected that the final target maximum number of copies does not exceed the copy number threshold, step F4 is immediately executed; if it is detected that the final target maximum number of copies exceeds the copy number threshold, step F10 is immediately executed.

[0333] Step F4: Check whether the final target maximum number of replicas is less than the latest number of Pods;

[0334] If it is detected that the final target maximum number of replicas is not less than the latest number of Pods, immediately execute step F5. If it is detected that the final target maximum number of replicas is less than the latest number of Pods, it means that after the Pod expansion is completed, the QPS tends to be stable, and the achieved value exceeds the maximum indicator threshold targetQps set by the target. The calculated perPodCpuRate and perPodMemoryRate also tend to a larger value, causing the final target maximum number of replicas to be smaller than the latest number of Pods. In this case, immediately execute step F6.

[0335] Step F5, replacing and updating the maximum number of HPA copies in the HPA object based on the final target maximum number of copies;

[0336] The update is completed by writing the final target maximum number of replicas selected by the scaling flag value into the maxReplicas field in the HPA object.

[0337] Step F6, obtaining the tag value of the shrinking-enabled tag;

[0338] The enableActiveScaleDown flag is a flag that controls the automatic scale down feature. It has the following functions:

[0339] When the tag value is set to true, the number of replicas can be automatically reduced (scaled down) based on resource utilization or preset rules.

[0340] If the tag value is set to false, scaling down is prohibited, and only scaling up or maintaining the current number of replicas is allowed.

[0341] The HPA PLUS controller obtains the current tag value of the flag by reading the custom enableActiveScaleDown tag in the HPA object.

[0342] Step F7, detecting whether the tag value is a preset value;

[0343] The default value is true as described above. If the label value is detected to be the preset value, that is, enableActiveScaleDown = true, active scaling can be triggered, so that the actual QPS approaches the target maximum indicator threshold targetQps, so that lossPods (lossPods = latest Pod number - final target maximum replica number) is close to 0, and step F8 needs to be executed. If the label value is detected to be not the preset value, that is, enableActiveScaleDown = false, active scaling is disabled, and step F9 needs to be executed.

[0344] Step F8, replacing and updating the maximum number of HPA copies in the HPA object based on the final target maximum number of copies;

[0345] Step F9: Replace and update the maximum number of HPA replicas in the HPA object based on the latest number of Pods;

[0346] Step F10: Replace and update the maximum number of HPA copies in the HPA object based on the copy number threshold.

[0347] When the final target maximum number of replicas exceeds the preset upper limit of the latest Pod number, the replica number threshold is updated to effectively avoid resource exhaustion caused by sudden traffic.

[0348] In the above technical solution, multi-level verification (threshold, Pod number comparison, label status) is used to ensure that the update operation is safe and reliable, effectively reducing the risk of update errors.

[0349] Following the previous point, the final target maximum number of replicas fluctuates significantly. This is mainly due to the calculated single pod CPU utilization and the maximum number of CPU replicas. There are three main reasons for these two large fluctuations:

[0350] 1. The QPS indicator data fluctuates;

[0351] 2. The current number of Pods fluctuates.

[0352] 3. Calculate the average CPU utilization (including soft interrupts) of the current instance and see fluctuations;

[0353] Currently, it is impossible to accurately estimate the QPS of an instance using a single model. Different types of requests have different QPS carrying capacities under unit resources.

[0354] Whether the connection between the client and the LB (Load Balancer) is long or short;

[0355] Whether the connection between LB and backend RS (Real Server) is long or short;

[0356] Whether the listener type is http (Hypertext Transfer Protocol) or https (Hypertext Transfer Protocol Secure);

[0357] TLS (Transport Layer Security) protocol version;

[0358] The specific cipher suite used by TLS;

[0359] Whether the TLS session uses the session / tickets session reuse mechanism, etc.

[0360] For all of the above reasons, we can only assume that the current distribution of request types is stable and then linearly infer the maximum number of Pods required.

[0361] To solve the above problem, it is necessary to smooth the CPU utilization of a single Pod and the maximum number of CPU replicas, and filter out abnormal values ​​with excessive deviation from the mean. The following processing method is used, with the CPU utilization of a single Pod and the maximum number of CPU replicas as the smoothing objects, and steps H1 to H5 are performed on each smoothing object:

[0362] Step H1, obtaining multiple historical data values ​​and current data values;

[0363] Keep at least five recently calculated historical data values ​​r5, r4, r3, r2, and r1 (sorted by time from earliest to latest) and the current data value rnow. If the smoothing object is single Pod CPU utilization, the above data values ​​are the historically calculated single Pod CPU utilization and the currently calculated single Pod CPU utilization;

[0364] If the smoothing object is the maximum number of CPU replicas, the above data values ​​are the historically calculated maximum number of CPU replicas and the currently calculated maximum number of CPU replicas.

[0365] Step H2, determining a mean and a standard deviation based on the plurality of historical data values ​​and the current data value;

[0366] Continuing with the previous example, the mean can be calculated using the following formula:

[0367]

[0368] The standard deviation can be calculated using the following formula:

[0369]

[0370] Step H3, determining whether the standard deviation-based noise filtering mechanism is satisfied based on the current data value, mean, and standard deviation;

[0371] The smooth object is detected as a noise point by using the standard deviation noise filtering mechanism. Indicates that the smoothed object is not a noise point, and step H4 is executed; if it is determined that the standard deviation noise filtering mechanism is not satisfied, that is, Indicates that the smoothed object is a noise point, and executes step H5.

[0372] Step H4, retain the current data value;

[0373] Step H5: retain the historical data value at the previous moment.

[0374] That is, the historical data value r1 at the previous moment is selected from the historical data values ​​r5, r4, r3, r2, and r1 for subsequent calculations.

[0375] In actual application, you can also set an expansion strategy to adjust the speed of expansion and contraction. Specifically, query the target expansion and contraction strategy attribute value corresponding to the first indicator data from the expansion and contraction strategy attribute value query table; use the target expansion and contraction strategy attribute value to update the contraction strategy attribute value in the HPA object.

[0376] The scaling strategy attribute value query table pre-stores a correspondence between the indicator data range of the first weight indicator and the scaling strategy attribute value, and the scaling strategy attribute value is used to characterize the scaling speed;

[0377] The above scaling policy attribute value query table may be a database table, an Excel spreadsheet, a configuration file, or other data structures, which are not limited here. The indicator data range is based on the targetQps value, as shown in Table 1 for ease of understanding:

[0378] Table 1

[0379] Indicator data range Scaling policy attribute value / second 0-0.3*targetQps 30 0.3*targetQps-0.5*targetQps 60 0.5*targetQps-0.7*targetQps 110 0.7*targetQps-0.85*targetQps 160 0.85*targetQps-0.95*targetQps 300

[0380] It should be noted that the correspondence between the indicator data range of the first weight indicator and the scaling strategy attribute value listed above is only an example. The specific correspondence between the indicator data range of the first weight indicator and the scaling strategy attribute value can be set according to actual needs and is not limited here.

[0381] Figure 7 The update process and calculation logic of the maximum number of replicas and the scaling policy attribute value in this embodiment are shown. The implementation of each module in the figure is consistent with the previous embodiment, and the specific details will not be repeated.

[0382] It should be noted that in actual application, Figure 7 The business flow in determines which business indicator is selected as the first weight indicator and the second weight indicator; Figure 7 The maximum quota in the value is the maximum metric threshold. When the first weighted metric is QPS, only targetQps needs to be set. When the first weighted metric is CPS, only targetCps needs to be set. The specific maximum metric threshold to be set depends on the instance specifications. a.cpuRatio and a.memoryRatio both need to be set in advance.

[0383] See also Figure 8 , is a block diagram of an embodiment of a device for dynamic resource scaling and quota restriction of a cloud load balancer provided in an embodiment of the present application. Figure 8 As shown, the device includes:

[0384] An acquisition module 800 is configured to periodically poll and obtain business data and resource usage used to calculate the current maximum number of Pod replicas, where the business data includes first indicator data of a first weight indicator and second indicator data of a second weight indicator; wherein the first weight indicator and the second weight indicator are both business indicators that do not require adaptation to the HPA controller;

[0385] A first determining module 801 is configured to determine a basic maximum number of replicas based on first indicator data and resource usage;

[0386] A second determination module 802 is configured to dynamically adjust the basic maximum number of replicas based on the second indicator data through a hierarchical threshold trigger mechanism to determine a final maximum number of replicas;

[0387] The updating module 803 is configured to update the HPA maximum number of copies in the HPA object using the final maximum number of copies.

[0388] This solution directly utilizes business metric data to determine the number of HPA replicas, completely avoiding the architectural limitations of traditional solutions that rely on adapters. Its core technological breakthrough lies in calculating the final maximum number of replicas based on the metric data of business metrics, such as the first and second weighted metrics, and dynamically updating the maximum number of HPA replicas in the HPA object. This entire process completely bypasses the intermediate adaptation step of converting business metrics into native HPA metrics, achieving seamless integration between business metrics and the elastic scaling system. This direct-connect design not only eliminates adapter development and maintenance costs, but also effectively avoids architectural redundancy, significantly reducing deployment configuration and version compatibility maintenance costs.

[0389] Furthermore, this technical solution optimizes the calculation of the maximum number of replicas through a hierarchical decision-making mechanism. It first generates a baseline maximum number of replicas based on resource utilization and a first weighted indicator, and then dynamically adjusts this value using a second weighted indicator. This multi-indicator collaborative scheduling design effectively overcomes the frequent pod number fluctuations caused by traditional independent triggering of each indicator, significantly improving scaling stability and further optimizing business performance.

[0390] Figure 9 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. Figure 9 The electronic device 1200 shown includes: at least one processor 1201, a memory 1202, at least one network interface 1204 and another user interface 1203. The various components in the electronic device 1200 are coupled together via a bus system 1205. It is understood that the bus system 1205 is used to achieve connection and communication between these components. In addition to including a data bus, the bus system 1205 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, the bus system 1205 is not described in detail. Figure 9 Various buses are labeled as bus system 1205.

[0391] The user interface 1203 may include a display, a keyboard, or a pointing device (eg, a mouse, a trackball, a touchpad, or a touch screen).

[0392] It is understood that the memory 1202 in the embodiment of the present application can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link DRAM (SLDRAM), and direct RAM bus random access memory (DRRAM). The memory 1202 described herein is intended to include, but is not limited to, these and any other suitable types of memory.

[0393] In some embodiments, the memory 1202 stores the following elements, executable units, or data structures, or a subset thereof, or an extended set thereof: an operating system 12021 and application programs 12022 .

[0394] Among them, the operating system 12021 includes various system programs, such as the framework layer, core library layer, and driver layer, which are used to implement various basic services and handle hardware-based tasks. Application 12022 includes various application programs, such as media players and browsers, which are used to implement various application services. The program that implements the method of the embodiment of the present application can be included in application 12022.

[0395] In an embodiment of the present application, by calling a program or instruction stored in the memory 1202, specifically, a program or instruction stored in the application 12022, the processor 1201 is used to execute the method steps provided in each method embodiment.

[0396] The method disclosed in the above embodiment of the present application can be applied to the processor 1201 or implemented by the processor 1201. The processor 1201 may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by the hardware integrated logic circuit or software instructions in the processor 1201. The above processor 1201 can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, or a discrete hardware component. The various methods, steps and logic block diagrams disclosed in the embodiments of the present application can be implemented or executed. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in conjunction with the embodiments of the present application can be directly embodied as being executed by a hardware decoding processor, or can be executed by a combination of hardware and software units in the decoding processor. The software unit can be located in a storage medium well-known in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in memory 1202. Processor 1201 reads information from memory 1202 and, in conjunction with its hardware, completes the steps of the above method.

[0397] It is understood that the embodiments described herein may be implemented using hardware, software, firmware, middleware, microcode, or a combination thereof. For hardware implementation, the processing unit may be implemented in one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSP devices, DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers, microprocessors, other electronic units for performing the functions described herein, or a combination thereof.

[0398] For software implementation, the technology described herein can be implemented by a unit that performs the functions described herein. The software code can be stored in a memory and executed by a processor. The memory can be implemented in the processor or outside the processor.

[0399] The electronic device provided in this embodiment may be Figure 9 The electronic device shown in FIG. 1 can perform the following operations: Figure 1-7 All steps of the method of dynamic resource scaling and quota limit of the cloud load balancer, thereby achieving Figure 1-7 The technical effects of the method of dynamic resource scaling and quota limit of the cloud load balancer shown in the figure are as follows. Figure 1-7 For the sake of brevity, the relevant description will not be repeated here.

[0400] The present application also provides a storage medium (computer-readable storage medium). The storage medium stores one or more programs. The storage medium may include volatile memory, such as random access memory; the memory may also include non-volatile memory, such as read-only memory, flash memory, hard disk, or solid-state drive; and the memory may also include a combination of the aforementioned types of memory.

[0401] When one or more programs in the storage medium can be executed by one or more processors, the above-mentioned method of dynamic resource scaling and quota restriction of the cloud load balancer can be implemented.

[0402] The processor is used to execute the program stored in the memory to implement the steps of the method for dynamic resource scaling and quota restriction of the cloud load balancer.

[0403] Professionals should also be further aware that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0404] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein may be implemented using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.

[0405] The specific implementation methods described above further illustrate the purpose, technical solutions and beneficial effects of this application. It should be understood that the above description is only the specific implementation methods of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of this application should be included in the scope of protection of this application.

Claims

1. A method for dynamic resource scaling and quota restriction of a cloud load balancer, characterized in that: The method comprises: Regularly polling to obtain business data and resource usage used to calculate the current maximum number of Pod replicas, the business data including first indicator data of a first weight indicator and second indicator data of a second weight indicator; wherein the first weight indicator and the second weight indicator are both business indicators that do not require adaptation to the HPA controller; Determine a basic maximum number of replicas based on the first indicator data and the resource usage rate; According to the second indicator data, dynamically adjusting the basic maximum number of replicas through a hierarchical threshold trigger mechanism to determine a final maximum number of replicas; The HPA maximum number of copies in the HPA object is updated using the final maximum number of copies.

2. The method according to claim 1, characterized in that The periodic polling obtains the business data and resource usage used to calculate the current maximum number of Pod replicas, including: Obtain a first resource usage sequence currently collected using a time window; Extracting the start and end timestamps of the time window; Determining whether the time span of the start and end timestamps is less than a preset time span; If it is determined that the time span of the start and end timestamps is less than the preset time span, querying the cluster monitoring system based on the preset time span to obtain the current first data sequence; if it is determined that the time span of the start and end timestamps is equal to or greater than the preset time span, querying the cluster monitoring system based on the time span of the start and end timestamps to obtain the current first data sequence; Time-aligning the first resource usage sequence and the first data sequence to determine a second resource usage sequence and the second data sequence; The resource usage rate is acquired based on the second resource usage rate sequence, and the first indicator data of the first weight indicator and the second indicator data of the second weight indicator are acquired based on the second data sequence.

3. The method according to claim 2, characterized in that The resource utilization includes the current number of Pods, current memory utilization, and current CPU utilization; The acquiring the resource usage rate based on the second resource usage rate sequence includes: Get the latest number of Pods from the second resource usage sequence; Determine the average number of Pods and the Pod standard deviation based on the second resource usage sequence; Determine whether a 2σ standard deviation noise filtering mechanism is satisfied based on the latest Pod number, the average Pod number, and the Pod standard deviation; If it is determined that the noise filtering mechanism based on the 2σ standard deviation is satisfied, the latest Pod number is determined as the current Pod number; If it is determined that the noise filtering mechanism based on the 2σ standard deviation is not satisfied, the latest Pod number at the previous running time is determined as the current Pod number; Acquire the latest memory utilization from the second resource utilization sequence, and determine the latest memory utilization as the current memory utilization; The latest CPU utilization is obtained from the CPU utilization sequence, and the latest CPU utilization is determined as the current CPU utilization.

4. The method according to claim 2, characterized in that acquiring first indicator data of the first weighted indicator and second indicator data of the second weighted indicator based on the second data sequence; Acquire the latest first indicator data from the second data sequence, and determine the latest first indicator data as the first indicator data of the first weighted indicator; The latest second indicator data is obtained from the second data sequence, and the latest second indicator data is determined as the second indicator data of the second weight indicator.

5. The method according to claim 3, characterized in that The determining of the basic maximum number of replicas based on the first indicator data and the resource usage rate includes: Obtaining a maximum indicator threshold, a memory utilization threshold, and a CPU utilization threshold of the first weighted indicator; Determine the maximum number of CPU replicas based on the current number of Pods, the maximum indicator threshold, the first indicator data, the CPU utilization threshold, and the current CPU utilization; Determine the maximum number of memory replicas based on the current number of Pods, the maximum indicator threshold, the first indicator data, the memory utilization threshold, and the current memory utilization; The maximum value is selected from the maximum number of CPU copies and the maximum number of memory copies, and the maximum value is rounded up to obtain the basic maximum number of copies.

6. The method according to claim 3, characterized in that The step of dynamically adjusting the basic maximum number of replicas based on the second indicator data through a hierarchical threshold triggering mechanism to determine a final maximum number of replicas includes: Determine a per-core unit value based on the second indicator data and the current number of Pods; Detecting whether the per-core unit value exceeds a preset indicator data threshold range; wherein the preset indicator data threshold range is a data range consisting of a preset pressure value and a limit value corresponding to each core; When it is detected that the per-core unit value is within the preset indicator data threshold range, obtaining a real-time change rate of the first weight indicator; detecting whether the real-time change rate exceeds a preset change rate threshold; When it is detected that the real-time change rate exceeds the preset change rate threshold, the basic maximum number of replicas is adjusted according to the preset adjustment number of Pods, and the adjusted basic maximum number of replicas is determined as the final maximum number of replicas; In the case where it is detected that the per-core unit value exceeds the preset indicator data threshold range, modifying the basic maximum number of replicas based on the second indicator data, and determining the basic maximum number of replicas after the modification as the final maximum number of replicas; When it is detected that the per-core unit value is lower than the preset indicator data threshold range, or when it is detected that the real-time change rate does not exceed the preset change rate threshold, the basic maximum number of copies is determined as the final maximum number of copies.

7. The method according to claim 6, characterized in that Modifying the basic maximum number of replicas based on the second indicator data includes: determining a desired number of cores based on the per-core unit value and the limit value; Determine the number of replicas to be added based on the required number of cores; The basic maximum number of replicas is modified based on the number of replica increases.

8. The method according to claim 5, characterized in that Before updating the HPA maximum number of replicas in the HPA object using the final maximum number of replicas, the method further includes: Determine a scaling flag value corresponding to the final maximum number of replicas based on the first indicator data, the current number of Pods, the current CPU utilization, the current memory utilization, the maximum indicator threshold, the CPU utilization threshold, the memory utilization threshold, and the maximum number of HPA replicas; The final maximum number of replicas and the corresponding scaling flag value are stored in a recommendation array. The recommendation array also includes the final maximum number of replicas determined at multiple historical moments and the corresponding scaling flag values ​​for each of the final maximum numbers of replicas.

9. The method according to claim 8, characterized in that The determining, based on the first indicator data, the current number of Pods, the current CPU utilization, the current memory utilization, the maximum indicator threshold, the CPU utilization threshold, the memory utilization threshold, and the maximum number of HPA replicas, a scaling flag value corresponding to the final maximum number of replicas includes: Determine a single Pod CPU utilization based on the first indicator data, the current number of Pods, and the current CPU utilization; and determine a single Pod memory utilization based on the first indicator data, the current number of Pods, and the current memory utilization; Determine a memory Pod incremental demand value based on the single Pod memory utilization, the first indicator data, and the maximum indicator threshold; and determine a CPU Pod incremental demand value based on the single Pod CPU utilization, the first indicator data, and the maximum indicator threshold; Determine the remaining available Pod capacity value of the CPU based on the CPU utilization threshold, the current CPU utilization, the current number of Pods, and the maximum number of HPA replicas; and determine the remaining available Pod capacity value of the memory based on the memory utilization threshold, the current memory utilization, the current number of Pods, and the maximum number of HPA replicas; The scaling flag value corresponding to the final maximum number of replicas is determined based on the final maximum number of replicas, the maximum number of HPA replicas, the memory Pod incremental requirement value, the CPU Pod incremental requirement value, the CPU remaining available Pod capacity value, the memory remaining available Pod capacity value, the first indicator data, the current CPU utilization rate, the current number of Pods, and the latest number of Pods.

10. The method according to claim 9, characterized in that The method further comprises: The single Pod CPU utilization and the maximum number of CPU replicas are respectively used as smoothing processing objects, and the following operations are performed on the smoothing processing objects: Get multiple historical data values ​​and current data values; determining a mean and a standard deviation based on the plurality of historical data values ​​and the current data value; Determining whether a standard deviation-based noise filtering mechanism is satisfied based on the current data value, the mean, and the standard deviation; If it is determined that the standard deviation noise filtering mechanism is satisfied, retaining the current data value; If it is determined that the standard deviation-based noise filtering mechanism is not satisfied, the historical data value at the previous moment is retained.

11. The method according to claim 9, characterized in that The determining of the scaling flag value corresponding to the final maximum number of replicas according to the final maximum number of replicas, the maximum number of HPA replicas, the memory Pod incremental requirement value, the CPU Pod incremental requirement value, the CPU remaining available Pod capacity value, the memory remaining available Pod capacity value, the first indicator data, the current CPU utilization rate, the current number of Pods, and the latest number of Pods includes: Check whether the final maximum number of replicas is the same as the latest number of Pods; When it is detected that the final maximum number of replicas is the same as the latest number of Pods, detect whether the final maximum number of replicas is selected from the maximum number of CPU replicas; When it is detected that the final maximum number of replicas is selected from the maximum number of CPU replicas, determining whether the first condition or the second condition is met based on the CPUPod increment requirement value and the CPU remaining available Pod capacity value; When the first condition or the second condition is met, setting the value of the telescopic flag to a first value; When it is detected that the final maximum number of replicas is selected from the maximum number of memory replicas, determining whether the third condition or the fourth condition is met based on the memory Pod incremental demand value and the memory remaining available Pod capacity value; When the third condition or the fourth condition is met, setting the value of the telescopic flag to the first value; When it is detected that the final maximum number of replicas is different from the latest number of Pods, determining the average data of the first indicator based on the second data sequence; Determining whether the fifth condition is satisfied based on the first indicator data and the first indicator average data, or determining whether the sixth condition is satisfied based on the current number of Pods and the current number of Pods at the previous running time; When the fifth condition or the sixth condition is met, setting the value of the telescopic flag to the second value; If the fifth condition and the sixth condition are not met, detecting whether the current CPU utilization is less than a preset pressure threshold; When detecting that the current CPU utilization is less than the preset pressure threshold, setting the scaling flag value to the third value; When it is detected that the current CPU utilization is not less than the preset pressure threshold, detecting whether the final maximum number of replicas is less than the latest number of Pods; When it is detected that the final maximum number of replicas is less than the latest number of Pods, the scaling flag value is set to the fourth value. When it is detected that the final maximum number of replicas is greater than the latest number of Pods, detecting whether the first indicator data is less than a preset initial threshold value; When detecting that the first indicator data is less than the preset initial threshold, setting the scaling flag value to the fifth value; When detecting that the first indicator data is not less than the preset initial threshold value, detecting whether the final maximum number of replicas is greater than the HPA maximum number of replicas; If it is detected that the final maximum number of replicas is not greater than the HPA maximum number of replicas, setting the scaling flag value to the sixth value; When it is detected that the final maximum number of replicas is greater than the HPA maximum number of replicas, the scaling flag value is set to the seventh value.

12. The method according to claim 8, characterized in that The updating of the HPA maximum number of copies in the HPA object using the final maximum number of copies includes: Traversing and detecting the scaling flag value corresponding to each of the final maximum number of replicas in the recommendation array; When it is detected that the scaling flag values ​​corresponding to the respective final maximum replica numbers in the recommended array are all greater than a preset value, or when there are values ​​that are both greater than and equal to the preset value, the smallest final maximum replica number is selected from the recommended array as the final target maximum replica number, and the final target maximum replica number is used to update the HPA maximum replica number in the HPA object; When it is detected that the scaling flag values ​​corresponding to the respective final maximum replica numbers in the recommended array are all less than the preset value, or when there are values ​​that are both less than and equal to the preset value, the largest final maximum replica number is selected from the recommended array as the final target maximum replica number, and the final target maximum replica number is used to update the HPA maximum replica number in the HPA object.

13. The method according to claim 12, characterized in that The updating of the HPA maximum number of copies in the HPA object using the final target maximum number of copies includes: Detecting whether the final target maximum number of replicas is the same as the HPA maximum number of replicas; When it is detected that the final target maximum number of replicas is the same as the HPA maximum number of replicas, not updating the HPA maximum number of replicas in the HPA object; If it is detected that the final target maximum number of replicas is different from the HPA maximum number of replicas, detecting whether the final target maximum number of replicas exceeds a replica number threshold; wherein the replica number threshold is a preset multiple upper limit of the latest Pod number; If it is detected that the final target maximum number of replicas does not exceed the replica number threshold, detecting whether the final target maximum number of replicas is less than the latest Pod number; When it is detected that the final target maximum number of replicas is not less than the latest number of Pods, the maximum number of HPA replicas in the HPA object is replaced and updated based on the final target maximum number of replicas; When it is detected that the final target maximum number of replicas is less than the latest number of Pods, a label value for enabling the scale-down label is obtained; Detecting whether the tag value is a preset value; When it is detected that the tag value is the preset value, replacing and updating the maximum number of HPA copies in the HPA object based on the final target maximum number of copies; When it is detected that the label value is not the preset value, the maximum number of HPA replicas in the HPA object is replaced and updated based on the latest number of Pods; When it is detected that the final target maximum number of replicas exceeds the replica number threshold, the HPA maximum number of replicas in the HPA object is replaced and updated based on the replica number threshold.

14. The method according to claim 1, wherein The method further comprises: Querying a target scaling policy attribute value corresponding to the first indicator data from a scaling policy attribute value query table; wherein the scaling policy attribute value query table pre-stores a correspondence between an indicator data range of the first weight indicator and a scaling policy attribute value, wherein the scaling policy attribute value is used to characterize a scaling speed; The scaling policy attribute value in the HPA object is updated using the target scaling policy attribute value.

15. A device for dynamic resource scaling and quota restriction of a cloud load balancer, characterized in that: The device comprises: An acquisition module, configured to periodically poll and obtain business data and resource usage used to calculate the current maximum number of Pod replicas, wherein the business data includes first indicator data of a first weight indicator and second indicator data of a second weight indicator; wherein the first weight indicator and the second weight indicator are both business indicators that do not require adaptation to the HPA controller; A first determining module, configured to determine a basic maximum number of replicas based on the first indicator data and the resource usage rate; A second determination module is configured to dynamically adjust the basic maximum number of replicas through a hierarchical threshold trigger mechanism based on the second indicator data to determine a final maximum number of replicas; An updating module is configured to update the maximum number of HPA copies in the HPA object using the final maximum number of copies.

16. An electronic device, characterized in that: include: A processor and a memory, wherein the processor is used to execute a program for dynamic resource scaling and quota restriction of a cloud load balancer stored in the memory, so as to implement the method for dynamic resource scaling and quota restriction of a cloud load balancer according to any one of claims 1 to 14.

17. A storage medium, characterized in that: The storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the method for dynamic resource scaling and quota restriction of a cloud load balancer according to any one of claims 1 to 14.

Citation Information

Patent Citations

  • Resource capacity expansion method and device, equipment and storage medium

    CN117149424A

  • Cloud collaboration container elastic expansion and contraction method and device and electronic equipment

    CN118041787A

  • Resource configuration method and device, electronic equipment and storage medium

    CN119718660A

  • Elastic storage method, device, electronic device and readable medium

    CN119759491A

  • Expansion method and device

    WO2020135633A1