Cloud service dynamic load balancing system based on AI intelligent agent

Through the dynamic load balancing system of cloud services based on AI agents, multi-dimensional data acquisition and dynamic model adjustment are used to solve the problem of unstable load distribution in high concurrency scenarios in cloud service systems, and more efficient resource allocation and system stability are achieved.

CN120301888BActive Publication Date: 2025-08-26SICHUAN ZHIXING ZHICHENG TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510780630.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-12
Publication Date
2025-08-26
Estimated Expiration
2045-06-12

AI Technical Summary

Technical Problem

The load distribution of existing cloud service systems is non-normal in high concurrency scenarios, and the static threshold is difficult to identify and instantaneous load surges, resulting in anomalies not being identified, affecting system stability and resource allocation efficiency.

Method used

The dynamic load balancing system of cloud services based on AI agents is adopted to build a service node's operating status image through multi-dimensional data collection, combine dynamic polynomial regression model and 3σ criterion to adjust the load prediction model parameters in real time, introduce traffic baseline multiples and resource occupancy calculation adjustment coefficients, and realize dynamic load balancing.

Benefits of technology

It significantly improves resource allocation efficiency and abnormal fuse accuracy in high concurrency scenarios, enhances the robustness and load balancing capabilities of the system, and reduces the impact of burst traffic on the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120301888B_ABST
    Figure CN120301888B_ABST
Patent Text Reader

Abstract

The present invention discloses a cloud service dynamic load balancing system based on AI intelligent body, which specifically relates to the field of dynamic load balancing technology. The method constructs a service node operation status portrait through multi-dimensional data collection, combines a dynamic polynomial regression model to realize nonlinear relationship modeling between RPS and load value, adopts a batch gradient descent algorithm with regularization parameters to iteratively optimize model parameters, and selects the optimal model through cross-validation. The system introduces a dual-threshold data queue management mechanism to ensure the timeliness and scale controllability of model training data, dynamically filters abnormal nodes based on the 3σ criterion, and designs an instantaneous traffic intensity perception module. The adjustment coefficient is calculated by comprehensively calculating the baseline multiple, resource occupancy rate and duration to realize dynamic correction of load prediction under burst traffic. Compared with traditional static strategies, the present invention significantly improves the resource allocation efficiency and overall system robustness in high-concurrency scenarios through the collaboration of AI model self-learning ability and real-time traffic perception.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of dynamic load balancing technology, and more specifically, to a cloud service dynamic load balancing system based on AI intelligent agents. Background Art

[0002] The widespread adoption of the internet and the expansion of its applications have made cloud computing platforms, as crucial infrastructure supporting internet applications, a current research hotspot. Cloud computing platforms provide convenient services to individuals and businesses in a virtualized, centralized manner across hardware, platform, and software layers. These services include virtual hosts, IP addresses, computing resources, development frameworks, monitoring services, productivity tools, application software, and a wide range of other internet-related services. Cloud computing platforms can be divided into two parts: the front-end, which provides the user access interface, and the back-end, which actually delivers the various services. From a back-end perspective, the various services of the cloud computing platform are provided by various server-side applications. The system composed of these server-side applications is called a cloud service system.

[0003] Performance issues in cloud service systems primarily stem from the continuous expansion of user bases during the development of the internet. The concurrent access of a large number of users inevitably creates immense concurrency pressure. Under high concurrency, any design issues in cloud service systems are magnified. Even millisecond-level latency can have a significant impact when faced with a large user base. Furthermore, the various concurrency issues that are difficult to reproduce and resolve under high concurrency also test the concurrency security of cloud service systems. Service downtime or data errors can have serious consequences for both businesses and users.

[0004] Existing technologies rely on a fixed mean and standard deviation calculated from historical data. However, cloud service loads are dynamic and time-varying, which can lead to non-normal load distribution. If the load deviates from the initial mean over a long period due to business growth, static thresholds can misclassify normal values ​​as abnormal. Instantaneous load surges caused by traffic bursts can be underestimated by static σ, leading to undetected anomalies.

[0005] In order to solve the above-mentioned defects, a technical solution is now provided. Summary of the Invention

[0006] In order to overcome the above-mentioned defects of the prior art, an embodiment of the present invention provides a cloud service dynamic load balancing system based on AI intelligent agent to solve the problems raised in the above-mentioned background technology.

[0007] To achieve the above object, the present invention provides the following technical solutions:

[0008] The cloud service dynamic load balancing system based on AI agent includes the following steps:

[0009] First, we collect service node operating metrics and establish a nonlinear relationship between RPS and load. We then normalize the data through feature scaling. We then iteratively train the model using a batch gradient descent algorithm with regularization parameters, and select the optimal model parameters based on cross-validation.

[0010] Combining the 3σ criterion and instantaneous traffic intensity perception to achieve intelligent scheduling, the three elements of traffic baseline multiple, resource occupancy rate, and duration are introduced to calculate the instantaneous traffic intensity level, and the dynamic adjustment coefficient is generated in combination with the load change rate to correct the load prediction model parameters in real time. .

[0011] In a preferred embodiment, the data collection includes a dynamic collection mechanism for multi-dimensional indicators, specifically including: periodic collection of four indicators: CPU usage, memory occupancy, number of requests per second, and response time; using a dual-channel redundant transmission protocol when reporting data to the central controller through a monitoring agent; using a time series database to store historical data, and achieving millisecond-level reading and writing of high-frequency indicators through memory cache.

[0012] In a preferred embodiment, the 3σ criterion dynamic filtering method calculates the load mean and standard deviation based on a moving time window; sets a dynamic abnormality threshold μ±3σ, and marks a node as an abnormal node when the predicted load value exceeds the threshold range; and activates an automatic circuit breaker mechanism for nodes that are marked three times in a row, and the circuit breaker time increases according to an exponential backoff algorithm.

[0013] In a preferred embodiment, real-time traffic is monitored by a traffic monitoring tool; when instantaneous traffic is generated, the instantaneous traffic duration and traffic intensity are recorded; and the instantaneous traffic intensity level is calculated based on the instantaneous traffic duration and traffic intensity.

[0014] In a preferred embodiment, the traffic intensity is determined by a combination of a baseline multiple and resource occupancy rate; the baseline multiple is based on the normal system traffic and is divided according to the excess ratio; the resource occupancy rate is the CPU / memory / bandwidth usage rate.

[0015] In a preferred embodiment, the load change rate is determined, and the adjustment coefficient is calculated by weighted summation in combination with the load change rate and the instantaneous traffic intensity level.

[0016] In a preferred embodiment, the adjustment coefficient is adjusted according to the value of the adjustment coefficient. The larger the adjustment coefficient, the greater the adjustment range.

[0017] In a preferred embodiment, the system includes the following modules: data collection and reporting module, model learning module, and dynamic request distribution module;

[0018] The data collection and reporting module is used to regularly collect the operating indicators of each service node and regularly report the data to the central controller through the monitoring agent; at the same time, it uses the time series database or memory cache to temporarily store data for model training;

[0019] The model learning module is used to store, process, and learn models after receiving data streams reported by the service cluster, and then write the models to the model storage. The model learning module consists of two main submodules: the data queue update module and the relationship model calculation module.

[0020] The dynamic request distribution module is used to predict the load value based on the model, filter abnormal nodes using the 3σ criterion, and achieve intelligent scheduling through a hybrid strategy.

[0021] In a preferred embodiment, the data queue update module is used to manage and update the data queue; the relational model calculation module is used to obtain the corresponding training model by dividing the data set, generating a feature matrix, feature scaling, training the candidate model, and cross-validation.

[0022] In a preferred embodiment, the dynamic request distribution module includes an adjustment module, which calculates the instantaneous flow intensity level according to the instantaneous flow duration and flow intensity, and then calculates the adjustment coefficient based on the load change rate and the instantaneous flow intensity level; The numerical value of .

[0023] The technical effects and advantages of the present invention are as follows:

[0024] The present invention constructs a service node operating status portrait through multi-dimensional data collection (CPU usage, memory occupancy, RPS, response time), combines a dynamic polynomial regression model to achieve nonlinear relationship modeling between RPS and load value, uses a batch gradient descent algorithm with regularization parameters to iteratively optimize model parameters, and selects the optimal model through cross-validation. The system introduces a dual-threshold data queue management mechanism to ensure the timeliness and scale controllability of model training data, dynamically filters abnormal nodes based on the 3σ criterion, and designs an instantaneous traffic intensity perception module. The adjustment coefficient is calculated by comprehensively calculating the baseline multiple, resource occupancy, and duration to achieve dynamic correction of load prediction under burst traffic. Compared with traditional static strategies, the present invention significantly improves the resource allocation efficiency, abnormal circuit breaking accuracy, and overall system robustness in high-concurrency scenarios through the collaboration of AI model self-learning capabilities and real-time traffic perception. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] In order to facilitate understanding by those skilled in the art, the present invention will be further described below with reference to the accompanying drawings;

[0026] Figure 1This is a flow chart of the cloud service dynamic load balancing system based on AI agent of the present invention;

[0027] Figure 2 Schematic diagram of the model learning process;

[0028] Figure 3 Schematic diagram of the calculation process for the relational model;

[0029] Figure 4 This is a structural diagram of the cloud service dynamic load balancing system based on AI intelligent agent of the present invention. DETAILED DESCRIPTION

[0030] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0031] The present invention constructs a service node operating status profile through multi-dimensional data collection (CPU usage, memory occupancy, RPS, response time), combines a dynamic polynomial regression model to achieve nonlinear relationship modeling between RPS and load value, uses a batch gradient descent algorithm with regularization parameters to iteratively optimize model parameters, and selects the optimal model through cross-validation. The system introduces a dual-threshold data queue management mechanism (minLen / maxLen) to ensure the timeliness and scale controllability of model training data, dynamically filters abnormal nodes based on the 3σ criterion, and designs an instantaneous traffic intensity perception module. The adjustment coefficient is calculated by comprehensively calculating the baseline multiple, resource occupancy, and duration to achieve dynamic correction of load prediction under burst traffic. Compared with traditional static strategies, the present invention significantly improves resource allocation efficiency, abnormal circuit breaking accuracy, and overall system robustness in high-concurrency scenarios through the synergy of AI model self-learning capabilities and real-time traffic perception.

[0032] Example 1

[0033] The present invention is based on the cloud service dynamic load balancing system of AI intelligent body, such as Figure 1 As shown, the following steps are included:

[0034] First, we collect the operating indicators of each service node and establish a nonlinear relationship between RPS and load value. We process the data through feature scaling and normalization, use the batch gradient descent algorithm with regularization parameters for iterative training, and select the optimal model parameters based on cross-validation. RPS stands for requests per second.

[0035] Intelligent scheduling is achieved by combining the 3σ criterion and instantaneous traffic intensity perception. The three elements of traffic baseline multiple, resource occupancy rate, and duration are introduced to calculate the instantaneous traffic intensity level. The dynamic adjustment coefficient is generated by combining the load change rate to correct the load prediction model parameters in real time.

[0036] Specifically,

[0037] Real-time collection of operational indicators of each service node, including CPU usage (Central Processing Unit), memory usage (measuring memory load), number of requests per second, and response time (quality of service indicator). Monitoring agents regularly report data to the central controller. A time series database or memory cache is used to temporarily store data for model training.

[0038] The model learning process is Figure 2 The model learning module receives the data stream reported by the service cluster, stores, processes and learns the data to obtain the model and writes the model to the model storage.

[0039] The model learning process can be simply described as follows: the model learning module receives the service node After reporting data, update first The data queue corresponds to the URL. The URL is the abbreviation of Uniform Resource Locator, which is an address standard used to uniquely identify and locate resources (such as web pages, images, files, etc.) on the Internet. Then, the model calculation is started according to the current model calculation status or the model calculation process is started after the current model calculation process is completed to ensure the synchronization of the model and the dataset.

[0040] The above brief description of the model learning process includes two main sub-processes: the data queue update process and the relational model calculation process. After the data queue update process completes the queue update, it initiates the relational model calculation process in a timely manner based on the model calculation status. The following describes the data queue update process and the relational model calculation process in detail.

[0041] Specifically, the data queue is updated; service nodes are defined The data queue is , using linked list as data structure, the queue length is , the queue element type is . Describes the model learning module receives Number of reports at a given moment Back pair The update process.

[0042] Specifically, it can be divided into three steps: adding data, adjusting queue length, and triggering the model calculation process at the right time:

[0043] Step 1: Add data

[0044] Step 1 is to convert the data Add to data queue . receive After that, the model learning module is updated according to different situations First, determine the data in Is there a data in , its RPS value and RPS value in If Exists, determine its load value Whether If they are equal, it means that the The data in has been learned, and the process ends at this time. Otherwise, the current and The load value corresponding to the RPS value is used replace in .if If it does not exist, then directly Add to End. This completes step 1. If not, proceed to step 2.

[0045] Step 2: Adjust the data queue length

[0046] Step 2 is to adjust the data queue Length In this strategy, There are two critical values, namely the minimum queue length minLen and the maximum queue length maxLen; minLen is the starting point for model calculation as a data set If the length is less than minLen, data needs to be accumulated to ensure the size of the data set required for model calculation and avoid underfitting of the model due to a small data set. maxLen is the maximum length of the data queue used as the data set for model calculation. If it exceeds maxLen, the queue needs to be pruned to ensure that the data queue does not grow indefinitely, limiting the size of the storage space.

[0047] After adding data, if If it is less than minLen, the model cannot be calculated at this time and the process ends; if If it is greater than maxLen, delete The oldest data in If the value is between minLen and maxLen, then go to step 3. At this point, step 2 is complete. If not, go to step 3.

[0048] Step 3: Trigger the model calculation process at the right time

[0049] Step 3 is to trigger the model calculation process in time according to the model calculation status. Indicates that in this strategy, X contains features represented by different powers of RPS values. Solution hypothesis In fact, the parameter vector θ is solved, so when storing the model, only the value of the parameter vector θ needs to be stored.

[0050] if Calculating, then mark After the calculation is completed, it is necessary to recalculate to ensure that the model information is consistent with the data set information. Otherwise, start the calculation , that is, start the relational model calculation process. In this way, the same There is only one calculation process at a time to avoid repeated calculations. At this point, the third stage is completed and the data queue update process ends.

[0051] The relational model calculation process realizes the conversion of data queue into model, that is, obtaining the hypothesis of the relationship between the RPS value and load value of the fitting service node. Figure 3 As shown in the figure, it includes five main steps: dividing the data set, generating feature matrix, feature scaling, training the candidate model, and cross-validation.

[0052] Step 1: Divide the dataset

[0053] Will The training set was randomly divided into two groups according to the ratio of 4:1. and cross validation set The random partitioning method makes the sample distribution in the training set and cross-validation set as consistent as possible with the sample distribution in the data queue, avoiding the impact of data set partitioning on model performance.

[0054] Step 2: Generate feature matrix and label matrix

[0055] The training set Convert to feature matrix and the label matrix ; Using polynomial functions to express assumptions ; ; where θ is the assumption The parameter vector of express The parameter of the jth item in , Fixed to 1; The value is RPS value , The value is The jth power of the polynomial function with different highest power d is different. Therefore, when the actual curve is uncertain, this strategy uses multiple d values ​​to calculate the model and selects the one with the best performance. At this time, the highest power d corresponds to To sum up, the assumptions in this strategy can be expressed as the following formula: ;

[0056] Feature-feature matrix and the label matrix The expression is as follows:

[0057] ; ;

[0058] In the above expression, m is equal to According to the above assumptions Description, accordingly, the characteristic matrix middle equal The jth power of the value of the i-th sample; the label matrix middle equal The i-th example in Value. Fill in according to the above rules and This completes step two.

[0059] Step 3: Feature Scaling

[0060] The third step is feature scaling, which is achieved by standardization. Execute the following formula in each column except the first one: ; Feature scaling can be achieved. In the above formula, represents the j-th column vector of the feature matrix X, express The mean of express The standard deviation of .

[0061] It should be noted that during the calculation process, the mean and standard deviation of each column except the first column need to be recorded, which is recorded as the mean vector and standard deviation vector . and Will be used when using the model to predict the load value of the service node, the mean vector It can be expressed by the following formula: ;in for The standard deviation vector can be expressed as follows: ;in for The standard deviation of the j+1th column in .

[0062] Step 4: Train the candidate model

[0063] Define the set of values ​​of the highest power d assumed as , for each item in the set , use the batch gradient descent algorithm after adding the regularization parameter to train the corresponding model The regularization parameter λ can be adjusted according to actual conditions.

[0064] Each one is described below The specific training process is as follows:

[0065] First, determine the initial θ value. Here, this paper designs an optimization scheme. If there is a previous And its highest power d is equal to If the same, use the last The value of θ in is used as the initial value. Otherwise, The +1-dimensional zero vector is used as the initial θ value.

[0066] Then, iterate and execute: ; The batch gradient descent formula including the regularization parameter is used until the stopping condition is met to obtain the value of θ. Finally, fill the value of θ into the formula: ; Get the current Next ,Right now .

[0067] Step 5: Cross-validation

[0068] First, generate a cross-validation set in the same way as the training set The feature matrix and label matrix of , and complete the feature scaling; then, for each of use The cost is calculated by the feature matrix and the label matrix; finally, the one with the minimum cost As .

[0069] Dynamic request distribution based on the 3σ principle: A dynamic request distribution strategy; real-time load prediction: Based on the current RPS input model, the theoretical load value of each node is predicted. Filtering through the 3σ principle mainly includes two steps: initialization and request distribution.

[0070] The initialization steps can be simply summarized as follows:

[0071] (1) Set a default load balancing strategy, such as a round-robin strategy.

[0072] (2) The request distribution module records the number of responses received from each service node in a fixed period and calculates the RPS value of each service node in the current period. For example, if the period is T, the number of responses received from the service node in the current period is The number of responses is N, and the current The RPS value is recorded as .

[0073] Here the request distribution module maintains each service node The role is to use the service node when requesting distribution , predict the load value of the service node, and then screen the service nodes according to the predicted load value, and finally optimize the load balancing effect.

[0074] The request distribution step is executed when the request reaches the load balancer. It can be simply summarized as follows:

[0075] (1) Request the distribution module to obtain the current service cluster set ;

[0076] (2) Initialize the service node set , minimum predicted load value ; Get each service node in P from the model storage in turn Model .if If it exists, then the mean vector recorded in the relational model learning process is and standard deviation vector Calculation interval . Mean vector and standard deviation vector Calculation relies on historical data, but cloud service load is dynamic and time-varying (such as sudden instantaneous traffic), which may cause the load distribution to be non-normal.

[0077] If the load deviates from the initial average value for a long time due to business growth, the static threshold will misjudge the normal value as abnormal. Underestimation leads to anomalies not being identified.

[0078] Monitor real-time traffic using traffic monitoring tools. When instantaneous traffic occurs, record its duration and intensity. Traffic intensity is determined by combining baseline multiples and resource utilization. Baseline multiples are based on the system's normal traffic flow and are divided by the percentage of excess traffic. Resource utilization is the percentage of CPU, memory, and bandwidth used.

[0079] The instantaneous flow intensity level is calculated according to the instantaneous flow duration and flow intensity. First, the instantaneous flow intensity level is normalized based on the instantaneous flow duration and flow intensity. Then, the instantaneous flow intensity level is calculated according to the following formula: Q=α*sj+γqd; where Q represents the instantaneous flow intensity level; sj represents the instantaneous flow duration, and qd represents the flow intensity; α and γ are the weight coefficients of the instantaneous flow duration and flow intensity, respectively.

[0080] Traffic intensity is determined by combining baseline multiples and resource utilization. The baseline multiple is based on the normal system traffic and is divided by the percentage of excess traffic. Resource utilization is the CPU / memory / bandwidth usage.

[0081] Furthermore, if the duration of the instantaneous traffic exceeds one second, only the instantaneous traffic intensity level within the moment of traffic surge is calculated.

[0082] The adjustment coefficient is calculated by combining the load change rate and the instantaneous flow intensity level. First, the load change rate and the instantaneous flow intensity level are normalized. Then, the adjustment coefficient is calculated using the following formula: T = β * fz + ξQ. Where T represents the adjustment coefficient, fz represents the load change rate, and β and ξ represent the weight coefficients of the load change rate and the instantaneous flow intensity level for calculating the adjustment coefficient.

[0083] Adjust according to the adjustment coefficient value The value of The value of is scaled accordingly. Dynamic calculation based on time window (shortened interval time) , adapt to the change of load trend. The value replaces the original value.

[0084] if , then use and Calculate current The predicted load value .if If it is less than minLoad, update Otherwise, join in .

[0085] (3) Use the default load balancing strategy from the collection Select the final service node .

[0086] In sub-step 2, the calculation formula is as follows: .in is the mean vector The j-th row element of is the standard deviation vector The j-th row element of for The j+1th row element of the parameter vector θ; n is the characteristic number, which is also equal to The highest power d assumed in . The actual eigenvalues ​​are scaled here because this strategy performs feature scaling in the relational model learning process. θ is obtained after feature scaling, and the corresponding actual eigenvalues ​​also need to be scaled before substituting them into the hypothesis.

[0087] In sub-step 3, this policy does not specify a specific default load balancing policy, but because The default load balancing strategy needs to be able to be selected when the service cluster changes dynamically. For example, the random strategy itself has this feature. In addition, for example, in the gateway load balancer in this article, if the default strategy adopts the polling strategy, it will skip the nodes that are not in the queue during the polling. Therefore, the default load balancing strategy can be designed based on the actual situation and the above characteristics.

[0088] Example 2

[0089] The present invention is based on the cloud service dynamic load balancing system of AI intelligent body, such as Figure 4 As shown, it includes the following modules: data collection and reporting module, model learning module, dynamic request distribution module;

[0090] The data collection and reporting module is used to regularly collect the operating indicators of each service node and regularly report the data to the central controller through the monitoring agent; at the same time, it uses the time series database or memory cache to temporarily store data for model training;

[0091] The model learning module is used to store, process, and learn models after receiving data streams reported by the service cluster, and then write the models to the model storage. The model learning module consists of two main submodules: the data queue update module and the relationship model calculation module.

[0092] The data queue update module is used to manage and update the data queue;

[0093] The relationship model calculation module is used to obtain the corresponding training model by dividing the data set, generating a feature matrix, feature scaling, training the candidate model, and cross-validation;

[0094] The dynamic request distribution module is used to predict load values ​​based on the model, filter abnormal nodes using the 3σ criterion, and implement intelligent scheduling through a hybrid strategy;

[0095] The dynamic request distribution module includes an adjustment module, which calculates the instantaneous flow intensity level according to the instantaneous flow duration and flow intensity, and then calculates the adjustment coefficient based on the load change rate and the instantaneous flow intensity level; The numerical value of .

[0096] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0097] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0098] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0099] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0100] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

Claims

1. Cloud service dynamic load balancing system based on AI agent, characterized by: The following steps are involved: First, we collect service node operating metrics and establish a nonlinear relationship between RPS and load. We then normalize the data through feature scaling. We then iteratively train the model using a batch gradient descent algorithm with regularization parameters, and select the optimal model parameters based on cross-validation. Combining the 3σ criterion and instantaneous traffic intensity perception to achieve intelligent scheduling, the three elements of traffic baseline multiple, resource occupancy rate, and duration are introduced to calculate the instantaneous traffic intensity level, and the dynamic adjustment coefficient is generated in combination with the load change rate to correct the load prediction model parameters in real time. ; Data collection includes a dynamic multi-dimensional indicator collection mechanism, specifically: periodically collecting four indicators: CPU usage, memory utilization, requests per second, and response time. A dual-channel redundant transmission protocol is used when reporting data to the central controller via the monitoring agent. A time series database is used to store historical data, and memory caching enables millisecond-level reading and writing of high-frequency indicators. The 3σ criterion dynamic filtering method calculates the load mean and standard deviation based on a moving time window; sets a dynamic abnormality threshold μ±3σ, and marks nodes as abnormal when the predicted load value exceeds the threshold range; and activates an automatic circuit breaker mechanism for nodes that are marked three times in a row, with the circuit breaker time increasing according to an exponential backoff algorithm; Monitor real-time traffic through traffic monitoring tools; when instantaneous traffic is generated, record the instantaneous traffic duration and traffic intensity; calculate the instantaneous traffic intensity level based on the instantaneous traffic duration and traffic intensity.

2. The AI ​​agent-based cloud service dynamic load balancing system according to claim 1, characterized in that: The traffic intensity is determined by the baseline multiple and resource utilization rate. The baseline multiple is based on the normal system traffic and is divided according to the excess ratio. The resource utilization rate is the CPU / memory / bandwidth usage.

3. The AI ​​agent-based cloud service dynamic load balancing system according to claim 1, characterized in that: Determine the load change rate, and calculate the adjustment coefficient by weighted summation based on the load change rate and the instantaneous flow intensity level.

4. The AI ​​agent-based cloud service dynamic load balancing system according to claim 3, characterized in that: Adjust according to the adjustment coefficient value The value of The values ​​are scaled accordingly.

5. The cloud service dynamic load balancing system based on AI agent according to claim 1 is characterized by: It includes the following modules: data collection and reporting module, model learning module, and dynamic request distribution module; The data collection and reporting module is used to regularly collect the operating indicators of each service node and regularly report the data to the central controller through the monitoring agent; at the same time, it uses the time series database or memory cache to temporarily store data for model training; The model learning module is used to store, process, learn and write models to the model storage after receiving the data stream reported by the service cluster; The model learning module consists of two main submodules: the data queue update module and the relational model calculation module; The dynamic request distribution module is used to predict the load value based on the model, filter abnormal nodes using the 3σ criterion, and achieve intelligent scheduling through a hybrid strategy.

6. The cloud service dynamic load balancing system based on AI agent according to claim 5 is characterized by: The data queue update module is used to manage and update the data queue; the relational model calculation module is used to obtain the corresponding training model by dividing the data set, generating the feature matrix, feature scaling, training the candidate model, and cross-validation.

7. The cloud service dynamic load balancing system based on AI agent according to claim 5 is characterized by: The dynamic request distribution module includes an adjustment module, which calculates the instantaneous flow intensity level according to the instantaneous flow duration and flow intensity, and then calculates the adjustment coefficient based on the load change rate and the instantaneous flow intensity level; The numerical value of .