Private cloud intelligent load balancing method, system and device and storage medium

By using the LSTM-Transformer hybrid model and load capacity index in private cloud load balancing for dynamic weight adjustment, combined with closed-loop feedback and online learning, the problems of insufficient dynamic adaptability and low prediction accuracy in the existing technology are solved, and more efficient resource utilization and lower maintenance costs are achieved.

CN120223700APending Publication Date: 2025-06-27ZHEJIANG ELECTRIC POWER DESIGN INST
View PDF 0 Cites 7 Cited by

Patent Information

Application Number
CN202510567427.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The existing private cloud load balancing technology has problems such as insufficient dynamic adaptability, low prediction accuracy, low resource utilization and high maintenance costs, especially in high concurrency scenarios, which are difficult to effectively respond to load changes.

Method used

The LSTM-Transformer hybrid model is used for load prediction, and the node weight is dynamically adjusted through the load capacity index (LCE), and combined with a closed-loop feedback system and online learning algorithm, the prediction model and traffic allocation strategy are optimized in real time.

Benefits of technology

It significantly improves the accuracy of load prediction, reduces prediction error and load variance, improves resource utilization, reduces maintenance costs, and achieves lower response time and timeout rates in high concurrency scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120223700A_ABST
    Figure CN120223700A_ABST
Patent Text Reader

Abstract

The invention discloses a private cloud intelligent load balancing method. The method comprises the following steps: S1, establishing an LSTM-Transform hybrid model to predict real-time load data of a private cloud node; s2, calculating a load capacity index according to the load prediction value and the current resource state of the node; s3, dynamically adjusting the weight value of each node based on the calculation result of the load capacity index; s4, sending the weight value to a load balancer, and adjusting a flow distribution strategy in real time; and S5, establishing a closed-loop feedback module for executing model parameter updating and dynamic threshold automatic adjustment operation according to the deviation between the actual load and the predicted value. According to the method, the dynamic weight adjustment mechanism based on the load capacity index is set, the model parameters are updated in real time in combination with the online learning algorithm, and the weight coefficient is optimized according to the index, so that the weight adjustment response delay is shortened from the minute level to the millisecond level, the load variance is reduced, and the resource utilization rate is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of resource management, and particularly to a private cloud intelligent load balancing method, system, device and storage medium. Background Art

[0002] In a private cloud environment, load balancing is a core technology to ensure high availability, resource utilization, stability and user experience in high-concurrency scenarios. Traditional methods mainly rely on the weighted round-robin (WRR) algorithm or a static weight allocation strategy based on thresholds, combined with a simple prediction model (such as LSTM) to achieve load adjustment.

[0003] However, the weighted round-robin (WRR) algorithm is prone to insufficient dynamic load adaptability, and static weight allocation is prone to load imbalance and resource waste. At the same time, due to the inability of traditional prediction models to capture complex load patterns, the prediction error is high. Therefore, the prediction model has defects in accuracy and real-time performance. For example, WRR dynamically adjusts node weights based on the current load, but relies on a preset threshold (such as CPU utilization > 80%) to trigger the adjustment, resulting in a lag in response and inability to handle sudden traffic; models such as LSTM can capture periodic load patterns, but the prediction error for mixed patterns (such as sudden traffic superimposed on periodic fluctuations) is as high as 20%-30% (MAPE), and there is a lack of real-time feedback mechanism and elastic expansion, and it is impossible to combine prediction with real-time data to achieve adaptive optimization.

[0004] In addition, cloud-native technologies (such as Kubernetes HPA) support elastic scaling, but rely on fixed thresholds to trigger, resulting in a delay in scaling and lack of coordination with weight adjustment, leading to resource waste (utilization rate is often lower than 70%).

[0005] In summary, the core defects of the existing technologies are reflected in three aspects: prediction accuracy, dynamic response ability and system integration efficiency. First, a single algorithm (such as LSTM) is difficult to capture the mixed pattern of long-term periodic load and sudden traffic at the same time, resulting in significant prediction deviation; second, weight adjustment and scaling strategies rely on static rules and cannot respond to sudden load changes (such as a sudden increase in requests) in real time, leading to node overload (CPU > 95%) or resource idleness (variance > 12%); finally, there is a lack of a closed-loop feedback mechanism, the model cannot online learn real-time data to optimize prediction, and the coordination between weight update and load balancer (such as Nginx) requires manual intervention, resulting in high maintenance costs.

[0006] Therefore, this application specifically proposes a private cloud intelligent load balancing method to solve the above technical problems. Summary of the Invention

[0007] The main purpose of the present invention is to provide a private cloud intelligent load balancing method to solve the technical problems proposed in the background art.

[0008] The present invention adopts the following technical solutions to solve the above technical problems:

[0009] A private cloud intelligent load balancing method, comprising the following steps:

[0010] S1. Establish an LSTM-Transformer hybrid model to predict the real-time load data of private cloud nodes;

[0011] S2. Calculate the load capacity index according to the load prediction value and the current resource status of the node;

[0012] S3. Dynamically adjust the weight values of each node based on the calculation result of the load capacity index;

[0013] S4. Send the weight value to the load balancer to adjust the traffic distribution strategy in real time;

[0014] S5. Establish a closed-loop feedback module for performing model parameter update and dynamic threshold automatic adjustment operations according to the deviation between the actual load and the prediction value.

[0015] Preferably, in the S1 step, the LSTM-Transformer hybrid model includes:

[0016] An LSTM module for capturing the long-term periodic characteristics of the load data, and the time window is set to a specified time period;

[0017] A Transformer module for identifying the global pattern characteristics of burst traffic, and the number of attention heads is set to a specified number;

[0018] A model output layer for fusing the prediction results of the LSTM module and the Transformer module with a specified weighting coefficient.

[0019] Preferably, the specific calculation formula for the load capacity index in the S2 step is:

[0020]

[0021] Where LCE i is the load capacity index of node i, C i is the remaining resource capacity of node i, L i is the real-time load of node i, L max is the maximum load threshold of node i, and k is a dynamic adjustment coefficient for controlling the steepness of the exponential term, and automatic adjustment is performed using an online learning algorithm according to the prediction error.

[0022] Preferably, the calculation formula for the weight value of each node in the S3 step is:

[0023]

[0024] Among them, W i is the weight value of node i, and LCE i is the load capacity index of node i. α is the weight adjustment coefficient used to control the influence of LCE on the weight, and β is the minimum weight threshold used to prevent the traffic allocation from failing due to too low node weights.

[0025] Preferably, the formula for calculating the deviation in step S5 is:

[0026]

[0027] Among them, MSE is the calculated deviation result, y t is used to represent the true load value at time t, is used to represent the load value predicted by the hybrid model at time t, and T is the number of data points within the prediction time window.

[0028] Preferably, the automatic adjustment operation in step S5 includes:

[0029] (1) Model parameter update: When the prediction error MSE exceeds the specified threshold, trigger the online learning algorithm to update the hybrid model parameters;

[0030] (2) Dynamic threshold adjustment: Automatically adjust the expansion trigger condition according to the historical load fluctuation amplitude;

[0031] (3) Parameter adaptive adjustment: If the load variance continues to be > 10%, increase α to enhance the influence of LCE. If the prediction error continues to be > 0.15, reduce the weighting coefficient of the Transformer module and increase the weight of the LSTM module.

[0032] Preferably, the closed-loop feedback module in step S5 includes:

[0033] An adaptive learning rate unit that dynamically adjusts the learning rate of the online learning algorithm according to the prediction error gradient;

[0034] A parameter constraint unit used to ensure that all automatically adjusted parameters are within the preset range;

[0035] A performance monitoring unit used to monitor the resource utilization rate, load variance, and request response time in real time, and trigger a comprehensive parameter optimization when any of the indicators exceeds the specified threshold.

[0036] Preferably, a private cloud intelligent load balancing system is connected to a load balancer and is used to execute any of the above-mentioned private cloud intelligent load balancing methods, including:

[0037] (a) An LSTM-Transformer hybrid model is used to predict the real-time load data of private cloud nodes, and combined with the current resource status of the nodes, the load capacity index is calculated;

[0038] (b) A closed-loop feedback module is used to calculate the deviation value between the actual load and the predicted value and perform model parameter update and dynamic threshold automatic adjustment operations

[0039] (c) An automatic parameter adjustment module includes: an online learning subunit for real-time updating of hybrid model parameters; a dynamic threshold generation subunit for automatically adjusting the expansion and weight adjustment thresholds according to load fluctuations; a parameter constraint subunit for ensuring that all adjusted parameters are within a preset range;

[0040] (d) A fallback mechanism module: used to restore to a stable parameter state when performance deteriorates.

[0041] In another aspect, the present invention also discloses a computer-readable storage medium storing a computer program, which when executed by a processor causes the processor to execute the steps of the above method.

[0042] In yet another aspect, the present invention also discloses a computer device including a memory and a processor, the memory storing a computer program, which when executed by the processor causes the processor to execute the steps of the above method.

[0043] As can be seen from the above technical solutions, the present invention provides a private cloud intelligent load balancing method and system. Compared with the prior art, the present invention has the following advantages:

[0044] 1. The present invention solves the core problems of long-term insufficient dynamic adaptability, low prediction accuracy, low resource utilization rate, and high maintenance cost in the field of private cloud load balancing through an innovative hybrid model architecture, dynamic parameter adaptive mechanism, closed-loop feedback system, and hardware-software co-design

[0045] 2. The present invention sets an LSTM-Transformer hybrid model. The LSTM module captures long-term periodic features through a long short-term memory network, and the Transformer module uses the self-attention mechanism to identify the global patterns of burst traffic, and balances the prediction results of the two through dynamic weighted fusion, reducing the prediction error and significantly improving the accuracy of load prediction, solving the defect that the traditional LSTM model cannot capture the mixed mode of periodic load and burst traffic at the same time.

[0046] 3. The present invention sets up a dynamic weight adjustment mechanism based on the load capacity index, combines an online learning algorithm to update the model parameters in real time, and automatically optimizes the weight coefficients according to indicators such as load variance and prediction error, shortening the weight adjustment response delay from the minute level to the millisecond level, reducing the load variance, and improving the resource utilization rate to solve the problem of the lag of dealing with static parameters.

[0047] 4. The present invention constructs a closed-loop feedback system to trigger the update of model parameters, dynamically adjust the expansion threshold and weight allocation strategy based on the prediction error, avoiding the expansion lag and resource waste caused by static thresholds in the traditional technology, and reducing the hardware cost.

[0048] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. Of course, any product implementing the present invention does not necessarily need to achieve all the above-mentioned advantages at the same time. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] The accompanying drawings forming a part of this application are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:

[0050] Figure 1 is a schematic diagram of the overall operation process of the method of the present invention;

[0051] Figure 2 is a schematic diagram of the data processing process of the system of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0052] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Without conflict, the embodiments in this application and the features in the embodiments can be combined with each other. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0053] In the embodiments, refer in detail to Figures 1 to 2 .

[0054] In the field of private cloud load balancing, existing technologies rely on static weight allocation strategies and single prediction models, resulting in problems such as insufficient dynamic adaptability, low prediction accuracy, and low resource utilization. The traditional Weighted Round Robin (WRR) algorithm adjusts traffic allocation through fixed weights or simple thresholds (such as CPU utilization thresholds), but it cannot respond to burst traffic in real time (such as the instantaneous load surge caused by a sudden increase in requests), resulting in node overload (such as CPU utilization exceeding 95%) or resource idleness (such as the utilization of other nodes being lower than 30%), with a load variance of over 12% and long-term resource utilization below 70%. At the same time, although existing prediction models (such as LSTM) can capture periodic load patterns, the prediction error for mixed patterns (such as burst traffic superimposed on periodic fluctuations) is as high as 20%-30%, and there is a lack of a closed-loop feedback mechanism, making it impossible to optimize model parameters or coordinate elastic scaling strategies online with real-time data, resulting in delayed expansion (such as only being triggered when the CPU continuously exceeds 90%), and the weights not being dynamically adjusted after new nodes are added, further exacerbating resource waste. In addition, the coordination between weight updates and load balancers (such as Nginx) relies on manual configuration or custom development, with high maintenance costs and poor scalability.

[0055] Therefore, an embodiment of the present invention proposes a private cloud intelligent load balancing method, which adopts a dynamic load balancing scheme based on an LSTM-Transformer hybrid model. By integrating the long-term dependency capture of LSTM for periodic loads and the global feature recognition of Transformer for burst traffic, the prediction error will be reduced, and the Load Capacity Index (LCE) will be introduced to achieve dynamic weight self-adaptation adjustment, improve resource utilization, and reduce load variance. At the same time, a closed-loop feedback system is constructed, integrating an online learning and elastic expansion trigger mechanism, to correct prediction deviations in real time and plan resources in advance, dynamically adjust the weights of new nodes to a balanced state, avoid delayed expansion and resource idleness, and achieve seamless docking with the load balancer through a standardized API interface, reducing maintenance costs. Finally, in high-concurrency scenarios, the response time is <1 second, the timeout rate is <5%, the stability is improved, and the hardware cost is reduced by 20%.

[0056] Specifically, as Figure 1 shown. A private cloud intelligent load balancing method based on machine learning prediction and LCE index proposed by an embodiment of the present invention includes the following steps:

[0057] S1. Establish an LSTM-Transformer hybrid model, and predict the real-time load data of private cloud nodes through the hybrid model.

[0058] Among them, the LSTM-Transformer hybrid model includes:

[0059] LSTM module, used to capture the long-term periodic characteristics of the load data, with the time window set to the past 12 to 24 hours;

[0060] Transformer module, used to identify the global pattern characteristics of the burst traffic, with the number of attention heads set to 4 to 8;

[0061] Model output layer, used to fuse the prediction results of the LSTM module and the Transformer module with a weighting coefficient of 0.4 to 0.6.

[0062] S2. Calculate the load capacity index based on the load prediction value and the current resource status of the node. At this time, the specific calculation formula of the load capacity index is:

[0063]

[0064] where, LCE i is the load capacity index of node i, C i is the remaining resource capacity of node i (such as the remaining number of CPU cores, the remaining memory capacity), L i is the real-time load of node i (such as the current CPU utilization rate, the number of requests), L max is the maximum load threshold of node i (such as the CPU upper limit is 100%), and k is a dynamic adjustment coefficient used to control the steepness of the exponential term, with an initial value of 0.1 to 1.0, and is automatically adjusted using an online learning algorithm according to the prediction error;

[0065] At this time, is used to calculate the proportion of the remaining resources of the node, reflecting the bearing potential of the node. At the same time, through e -k(Lmax-Li) , when the node load is close to the threshold (L i ≈L max ), the exponential term approaches 0, and the LCE decreases significantly, suppressing the weight allocation; conversely, the LCE of the low-load node increases and the weight increases.

[0066] Balance: Combine the remaining resources and the load trend to avoid overloading or idling.

[0067] where, k is the dynamic adjustment coefficient, with an initial value of 0.1 to 1.0, and is automatically adjusted according to the prediction error through an online learning algorithm.

[0068] S3. Based on the calculation result of the load capacity index, dynamically adjust the weight values of each node. At this time, the calculation formula of the weight value of each node is:

[0069]

[0070] where, W i is the weight value of node i, LCEi is the load capacity index of node i, α is the weight adjustment coefficient (ranging from 0.5 to 1.0) used to control the influence of LCE on the weight, and β is the minimum weight threshold (ranging from 0.0 to 0.2) used to prevent the traffic allocation from failing due to too low node weights;

[0071] At this time, is used to normalize LCE to the base value of the weight, amplify or reduce the influence of LCE through α (such as increasing α to 0.8 - 1.0 during burst traffic), and β ensures that cold nodes retain the base weight for dynamic parameter adjustment.

[0072] S4. Send the weight value to the load balancer to adjust the traffic allocation policy in real time, and the weight update period is 1 to 5 seconds;

[0073] S5. Establish a closed-loop feedback module for performing model parameter update and dynamic threshold automatic adjustment operations according to the deviation between the actual load and the predicted value. At this time, the calculation formula for the deviation is:

[0074]

[0075] where MSE is the calculated deviation result, y t is used to represent the true load value at time t (such as CPU utilization rate), is used to represent the load value predicted by the hybrid model at time t, and T is the number of data points within the prediction time window (such as 60 data points at 1-minute intervals);

[0076] And the automatic adjustment operation generally includes the weight adjustment coefficient α and the minimum weight threshold β being automatically optimized according to the load variance through the closed-loop feedback module, specifically including:

[0077] (1) Model parameter update: When the prediction error MSE exceeds the threshold of 0.05 to 0.2, trigger the online learning algorithm to update the hybrid model parameters;

[0078] (2) Dynamic threshold adjustment: Automatically adjust the capacity expansion trigger condition according to the historical load fluctuation range (such as the CPU utilization rate threshold is dynamically adjusted from 85% to 90% to 80% to 88%);

[0079] (3) Parameter adaptive adjustment: If the load variance continues to be > 10%, increase α (to 0.8 to 1.0) to enhance the influence of LCE. If the prediction error continues to be > 0.15, reduce the weighting coefficient of the Transformer module to 0.3 to 0.5 and increase the weight of the LSTM module.

[0080] Furthermore, the closed-loop feedback module includes:

[0081] An adaptive learning rate unit that dynamically adjusts the learning rate of the online learning algorithm according to the prediction error gradient (e.g., the initial learning rate is from 0.001 to 0.01, and when the error increases, the learning rate is increased to 0.005 to 0.05);

[0082] A parameter constraint unit that ensures that all automatically adjusted parameters are within a preset range (e.g., k ∈ [0.1, 1.0], α ∈ [0.5, 1.0]);

[0083] A performance monitoring unit that is used to monitor the resource utilization rate, load variance, and request response time in real time, and triggers a comprehensive parameter optimization when any of the metrics exceeds a specified threshold (e.g., the resource utilization rate < 70% or the response time > 1 second).

[0084] At this time, the automatic adjustment mechanism includes the following triggering conditions:

[0085] (a) Model parameter update triggering conditions:

[0086] The prediction error (MSE) exceeds the threshold of 0.1 to 0.2 for 3 to 5 consecutive times;

[0087] Or the real-time load fluctuation amplitude exceeds 1.5 to 2 times the standard deviation of the historical mean;

[0088] (b) Weight coefficient adjustment triggering conditions:

[0089] The load variance exceeds the threshold of 8% to 12%;

[0090] Or the difference in resource utilization rate between nodes > 20%;

[0091] (c) Dynamic threshold adjustment triggering conditions:

[0092] When the predicted load continuously exceeds the expansion threshold of 85% to 90% and no expansion is triggered, the expansion threshold is automatically reduced to 80% to 85%;

[0093] Or when the utilization rate of the node after expansion < 30%, the expansion threshold is automatically increased to 90% to 95%.

[0094] In addition, the above automatic adjustment mechanism adopts the following optimization strategies:

[0095] (a) Gradient descent optimization: Minimize the prediction error (MSE) through the backpropagation algorithm and update the parameters of the LSTM and Transformer modules;

[0096] (b) Adaptive weight allocation: According to the load type (e.g., when the proportion of burst traffic > 50%, increase the weighted coefficient of the Transformer module to 0.6 to 0.8);

[0097] (c) Parameter annealing mechanism: When the prediction error continues to decrease, gradually reduce the adjustment step size of k (e.g., from 0.1 to 0.05).

[0098] Furthermore, it should be further noted that the automatic adjustment mechanism realizes parameter fallback in the following way: When the performance metrics (such as resource utilization or response time) deteriorate after automatic adjustment, trigger the parameter to fallback to the nearest stable state, or adopt preset default parameters (such as k = 0.5, α = 0.7).

[0099] On the other hand, the present invention also discloses a private cloud intelligent load balancing system, referring to Figure 2 , connected to the load balancer, for executing the private cloud intelligent load balancing method in the above embodiments, including:

[0100] (a) LSTM-Transformer hybrid model, adopting a dual-branch parallel structure, for predicting the real-time load data of private cloud nodes, and combining the current resource status of the nodes to calculate the load capacity index;

[0101] (b) Closed-loop feedback module, for calculating the deviation value between the actual load and the predicted value and performing model parameter update and dynamic threshold automatic adjustment operations

[0102] (c) Automatic parameter adjustment module, including: an online learning subunit, for real-time updating of the hybrid model parameters; a dynamic threshold generation subunit, for automatically adjusting the expansion and weight adjustment thresholds according to load fluctuations; a parameter constraint subunit, for ensuring that all adjusted parameters are within the preset range;

[0103] (d) Fallback mechanism module: for restoring to the stable parameter state when the performance deteriorates.

[0104] In summary, through an innovative hybrid model architecture, dynamic parameter adaptive mechanism, closed-loop feedback system, and hardware-software co-design, the method and system solve the core problems that have long existed in the field of private cloud load balancing, such as insufficient dynamic adaptability, low prediction accuracy, low resource utilization, and high maintenance costs. First, aiming at the defect that the traditional LSTM model cannot capture the mixed mode of periodic load and burst traffic at the same time, the present invention proposes an LSTM-Transformer hybrid model: the LSTM module captures long-term periodic features (such as daily traffic peaks) through long short-term memory networks, and the Transformer module uses the self-attention mechanism to identify the global patterns of burst traffic (such as instantaneous request surges), and balances the prediction results of the two through dynamic weighted fusion (weight range 0.4-0.6), reducing the prediction error (MAPE) from 20%-30% of traditional technologies to 9.3%, significantly improving the accuracy of load prediction. Second, to address the lag of static parameters, the present invention designs a dynamic weight adjustment mechanism based on the load capacity index (LCE), combines an online learning algorithm (such as the Adam optimizer) to update the model parameters in real time, and automatically optimizes the weight coefficients (such as α increasing from 0.5 to 0.8) according to indicators such as load variance and prediction error, shortening the weight adjustment response delay from the minute level to the millisecond level, reducing the load variance from 12% to 8.2%, and increasing the resource utilization rate to 85%. Further, the closed-loop feedback system triggers model parameter updates, dynamically adjusts the expansion threshold (such as the CPU utilization rate decreasing from 90% to 85%) and weight allocation strategies through prediction errors, avoiding the expansion lag and resource waste caused by static thresholds in traditional technologies, and reducing the hardware cost by 20%. In addition, the hardware-software co-design of the standardized API interface (such as gRPC) and the distributed architecture (central server + edge nodes) realizes the seamless docking of the load balancer (such as Nginx) and the prediction model, reduces the manual configuration requirements, reduces the maintenance cost by 30%, and supports rapid expansion to multi-business scenarios. Finally, through the comprehensive optimization of prediction accuracy, dynamic response, resource utilization, and system scalability, the present invention achieves the technical effects of a request timeout rate <5% and a 40% improvement in system stability in high-concurrency scenarios, providing an intelligent and adaptive solution for private cloud load balancing.

[0105] In a specific embodiment, a central prediction server (responsible for global model training and parameter synchronization) and edge computing nodes (5) are used to perform corresponding operations, and there are:

[0106] (1) Parameter configuration:

[0107] Hybrid model parameters: LSTM module time window: 12 hours (capturing the daily periodic pattern of design tasks, such as the surge in rendering tasks at 3 pm); Transformer number of attention heads: 8 (identifying burst tasks, such as when users submit multiple complex model renderings simultaneously); Dynamic weight coefficient: LSTM weight 0.5, Transformer weight 0.5 (balancing periodic and burst task predictions).

[0108] Load Capacity Index (LCE) parameters: k = 0.7 (strengthening the suppression of critical load nodes to avoid GPU idleness and CPU overload); Scaling trigger threshold: CPU utilization > 85% or GPU utilization > 90% and remaining resources < 20%.

[0109] Automatic adjustment rules: When the load variance > 10%, increase the weight adjustment coefficient α to 0.9; If the prediction error (MSE) is > 0.1 for 3 consecutive times, reduce the Transformer weight to 0.4 (giving priority to ensuring periodic tasks).

[0110] (2) Implementation steps and process

[0111] First step, perform real-time data collection and preprocessing operations:

[0112] The data source is node load data, specifically including: collecting CPU, GPU utilization, memory bandwidth, and network latency per second;

[0113] Business characteristics include: task type (modeling / rendering / simulation), user concurrency, task priority (such as marking rendering tasks as "high priority").

[0114] The specific operation process of feature input operations includes: The LSTM module inputs the load time series of the past 12 hours (such as the historical fluctuations of GPU utilization); The Transformer module inputs the real-time task queue length (such as the rendering task queue suddenly increasing to 100) and the cross-node GPU correlation matrix (correlation coefficient 0.85).

[0115] Second step, perform hybrid model prediction and weight allocation operations:

[0116] Use the LSTM module to predict the periodic load in the next 10 minutes: Predict the average CPU utilization of 80% and the average GPU utilization of 75% (based on the historical task peak at 3 pm);

[0117] Then use the Transformer module to identify burst tasks: Detect that users submit 20 high-precision rendering tasks simultaneously, and predict that the GPU utilization peak reaches 95%;

[0118] Finally, perform weighted fusion: Integrate the prediction results of LSTM and Transformer. The final predicted GPU utilization rate is 92%, triggering dynamic weight adjustment.

[0119] In the third step, perform LCE calculation and weight assignment operations, as illustrated by the following data examples:

[0120] Node A (GPU utilization rate: 90%): LCE = 0.25 (insufficient remaining resources, weight reduced);

[0121] Node B (GPU utilization rate: 60%): LCE = 0.4 (weight increased to 0.35);

[0122] Node C (GPU idle): LCE = 0.5 (weight increased to 0.4, preferentially allocate burst rendering tasks).

[0123] In the fourth step, perform closed-loop feedback and automatic adjustment operations:

[0124] (a) Prediction error trigger: The actual GPU peak reaches 98%, MSE = 0.12 (exceeding the threshold of 0.1). Therefore, trigger model parameter update, adjust the attention weights of Transformer through the Adam optimizer, and enhance the ability to identify high-priority tasks;

[0125] (b) Dynamic expansion trigger: The GPU utilization rate of Node A continuously > 90%. Therefore, trigger expansion, add Node D (initial weight set to 70% of the current average weight), and then perform weight self-adaptation adjustment operations: The LCE of Node D is 0.3, and the weight is gradually increased to 0.28, and the load variance is reduced from 15% to 8%.

[0126] (c) Resource scheduling optimization: High-priority rendering tasks are preferentially allocated to GPU idle nodes, and modeling tasks are scheduled to nodes with sufficient remaining CPU resources.

[0127] (3) Implementation effects

[0128] Resource utilization improvement: The average GPU utilization rate is increased from 75% to 88%, the CPU utilization rate is increased from 60% to 78%, and the hardware cost is reduced by 25%;

[0129] Task response speed: The waiting time for burst rendering tasks is shortened from 12 minutes to 3 minutes, and the timeout rate is reduced from 15% to 5%;

[0130] Enhanced stability: Under 100 concurrent rendering tasks, the system does not experience node overload (both CPU / GPU < 95%), and the task completion time fluctuation < 5%.

[0131] In another aspect, the present invention also discloses a computer-readable storage medium storing a computer program, which when executed by a processor causes the processor to execute the steps of the above method.

[0132] In yet another aspect, the present invention also discloses a computer device including a memory and a processor, where the memory stores a computer program, which when executed by the processor causes the processor to execute the steps of the above method.

[0133] In yet another embodiment provided by the present application, there is also provided a computer program product containing instructions, which when run on a computer causes the computer to execute any of the private cloud intelligent load balancing methods in the above embodiments.

[0134] It can be understood that the system provided by the embodiments of the present invention corresponds to the method provided by the embodiments of the present invention. For the explanations, examples, and beneficial effects of related content, reference can be made to the corresponding parts in the above method.

[0135] The embodiments of the present application also provide an electronic device, including a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete communication with each other through the communication bus.

[0136] The memory is used to store a computer program.

[0137] The processor is used to implement the above private cloud intelligent load balancing method when executing the program stored on the memory.

[0138] The communication bus mentioned in the above electronic device may be a peripheral component interconnect standard bus or an extended industry standard architecture bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc.

[0139] The communication interface is used for communication between the above electronic device and other devices.

[0140] The memory may include a random access memory and may also include a non-volatile memory, such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.

[0141] The above processor may be a general-purpose processor, including a central processing unit, a network processor, etc.; it may also be a digital signal processor, an application-specific integrated circuit, a field programmable gate array, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0142] It should also be noted that the electronic device further includes a terminal device, which can also be referred to as a terminal, user equipment, mobile station, mobile terminal, etc. The terminal device can be a mobile phone, smart TV, wearable device, tablet computer, computer with wireless transceiver function, virtual reality terminal device, augmented reality terminal device, wireless terminal in industrial control, wireless terminal in driverless, wireless terminal in remote surgery, wireless terminal in smart grid, wireless terminal in transportation safety, wireless terminal in smart city, wireless terminal in smart home, and so on. The embodiments of the present application do not limit the specific technologies and specific device forms adopted by the terminal device.

[0143] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server, data center, etc. that contains one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium, or a semiconductor medium (such as a solid-state drive), etc.

[0144] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included in the protection scope of the present invention.

[0145] In addition, it should be noted that if there are directional indications (such as up, down, left, right, front, back...) involved in the embodiments of the present invention, the directional indications are only used to explain the relative positional relationship and movement conditions between components in a specific posture. If the specific posture changes, the directional indications will also change accordingly.

[0146] In addition, if there are descriptions involving "first", "second", etc. in the embodiments of the present invention, such descriptions of "first", "second", etc. are for descriptive purposes only and should not be construed as indicating or implying their relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In addition, the meaning of "and / or" appearing throughout the text includes three parallel scenarios. Taking "A and / or B" as an example, it includes Scenario A, or Scenario B, or the scenario where both A and B are satisfied simultaneously. In addition, in the embodiments of the present invention, "a plurality of" means more than two. In addition, the technical solutions between various embodiments can be combined with each other, but it must be based on the ability of those of ordinary skill in the art to implement. When the combination of technical solutions results in contradictions or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection required by the present invention.

Claims

1. A private cloud intelligent load balancing method, characterized in that: The following steps are involved: S1. Establish an LSTM-Transformer hybrid model and use the hybrid model to predict the real-time load data of private cloud nodes; S2. Calculate the load capacity index based on the load prediction value and the current resource status of the node; S3. Dynamically adjust the weight value of each node based on the calculation result of the load capacity index; S4. Send the weight value to the load balancer to adjust the traffic distribution strategy in real time; S5. Establish a closed-loop feedback module to perform model parameter update and dynamic threshold automatic adjustment operations according to the deviation between the actual load and the predicted value.

2. The private cloud intelligent load balancing method according to claim 1, characterized in that: The LSTM-Transformer hybrid model in step S1 includes: LSTM module, used to capture the long-term periodic characteristics of load data, and the time window is set to a specified time period; Transformer module, used to identify the global pattern characteristics of burst traffic, with the number of attention heads set to a specified number; The model output layer is used to fuse the prediction results of the LSTM module and the Transformer module with specified weight coefficients.

3. The private cloud intelligent load balancing method according to claim 1, characterized in that: The specific calculation formula of the load capacity index in step S2 is: Among them, LCE i is the load capacity index of node i, C i is the remaining resource capacity of node i, L i is the real-time load of node i, L max is the maximum load threshold of node i, k is the dynamic adjustment coefficient used to control the steepness of the exponential term, and automatic adjustment is performed using an online learning algorithm based on the prediction error.

4. The private cloud intelligent load balancing method according to claim 1, characterized in that: The weight value calculation formula of each node in step S3 is: Among them, W i is the weight value of node i, LCE i is the load capacity index of node i, α is the weight adjustment coefficient, which is used to control the influence of LCE on the weight, and β is the minimum weight threshold, which is used to prevent the node weight from being too low and causing traffic distribution failure.

5. The private cloud intelligent load balancing method according to claim 4, characterized in that: The calculation formula of the deviation in step S5 is: Among them, MSE is the deviation result obtained by calculation, y t Used to represent the actual load value at time t, It is used to represent the load value predicted by the hybrid model at time t, where T is the number of data points in the prediction time window.

6. The private cloud intelligent load balancing method according to claim 5, characterized in that: The automatic adjustment operation in step S5 includes: (1) Model parameter update: When the prediction error MSE exceeds the specified threshold, the online learning algorithm is triggered to update the hybrid model parameters; (2) Dynamic threshold adjustment: Automatically adjust the expansion trigger conditions based on historical load fluctuations; (3) Parameter adaptive adjustment: If the load variance is continuously >10%, α is increased to enhance the impact of LCE. If the prediction error is continuously >0.15, the weight coefficient of the Transformer module is reduced and the weight of the LSTM module is increased.

7. The private cloud intelligent load balancing method according to claim 1, characterized in that: The closed-loop feedback module in step S5 includes: Adaptive learning rate unit, which dynamically adjusts the learning rate of the online learning algorithm according to the prediction error gradient; A parameter constraint unit, used to ensure that all automatically adjusted parameters are within a preset range; The performance monitoring unit is used to monitor resource utilization, load variance, and request response time in real time, and trigger comprehensive parameter optimization when any indicator exceeds the specified threshold.

8. A private cloud intelligent load balancing system, connected to a load balancer, for executing the private cloud intelligent load balancing method according to any one of claims 1 to 7, characterized in that: include: (a) LSTM-Transformer hybrid model, which is used to predict the real-time load data of private cloud nodes and calculate the load capacity index based on the current resource status of the nodes; (b) Closed-loop feedback module, used to calculate the deviation between the actual load and the predicted value and perform model parameter updates and dynamic threshold automatic adjustment operations (c) Automatic parameter adjustment module, including: an online learning subunit for updating hybrid model parameters in real time; a dynamic threshold generation subunit for automatically adjusting expansion and weight adjustment thresholds according to load fluctuations; and a parameter constraint subunit for ensuring that all adjustment parameters are within a preset range; (d) Fallback mechanism module: used to restore to a stable parameter state when performance deteriorates.

9. A computer-readable storage medium, characterized in that: A computer program is stored, and when the computer program is executed by a processor, the processor is caused to perform the steps of the method according to any one of claims 1 to 7.

10. A computer device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Load adjustment method and device for container instance and readable storage medium

    CN120560785A

  • A method, device and readable storage medium for load adjustment of a container instance

    CN120560785B

  • Request processing method and device, medium and program product

    CN120692100A

  • Industrial control data processing system and method based on embedded real-time operating system

    CN120762376A

  • Streaming media video sharing platform cluster management data processing method and system

    CN121078242A