Load balancing method and device, electronic equipment and storage medium
By collecting historical server load data and benchmark response time, using the CNN-Bi-LSTM-Attention model to predict future load, calculating comprehensive evaluation values and distributing requests, this solves the lack of forward-looking load balancing in existing technologies and achieves precise scheduling and resource optimization for burst traffic.
Patent Information
- Application Number
- CN202511248036.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-02
- Publication Date
- 2025-10-17
AI Technical Summary
Existing load balancing solutions lack foresight, resulting in the system being unable to cope with sudden traffic and achieve accurate scheduling optimization.
By collecting historical multi-dimensional load data and benchmark response time of the server cluster, the weight coefficient of the load indicator is determined, and the CNN-Bi-LSTM-Attention hybrid deep learning model is used to predict future load and response time. The comprehensive load evaluation value is calculated, and requests are distributed based on the evaluation value.
It enables a preliminary assessment of the server's future load-bearing pressure, and can complete precise scheduling before burst traffic arrives, thereby improving system throughput and resource utilization and avoiding node overload.
Smart Images

Figure CN120803739A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of load balancing, and particularly relates to a load balancing method and device, electronic equipment and a storage medium. BACKGROUND
[0002] With the continuous expansion of Internet business scale, distributed server clusters have become the infrastructure to support various online services. Load balancing technology, as a core component of distributed systems, plays a role in reasonably distributing network requests to multiple servers to improve the overall throughput, availability and resource utilization of the system, and to avoid single node overload. Traditional load balancing strategies mainly make decisions based on static rules or simple real-time indicators, which are difficult to adapt to complex and variable modern application scenarios with different business characteristics.
[0003] Existing load balancing schemes usually make decisions based on static rules or simple real-time indicators. For example, the round robin method assigns requests to each server in a fixed order, and the least connection method preferentially assigns new requests to the node with the fewest current connections. More advanced dynamic schemes collect real-time performance indicators of server nodes (such as CPU (Central Processing Unit) usage, memory occupancy), and calculate weights based on these instantaneous data, and use weighted round robin and other methods for traffic distribution.
[0004] However, the above schemes all rely on the current or past system state for judgment, and the decision mechanism lacks foresight, resulting in the system being unable to cope with sudden traffic and achieving precise scheduling optimization. SUMMARY
[0005] The present application provides a load balancing method, device, electronic equipment and storage medium to solve the problem that the decision mechanism lacks foresight in the prior art, resulting in the system being unable to cope with sudden traffic and achieving precise scheduling optimization.
[0006] In a first aspect, the present application provides a load balancing method, comprising:
[0007] Collecting multi-dimensional load data and baseline response time of each server in a historical time period in a server cluster; determining index weight coefficients of each load indicator in the multi-dimensional load data according to the multi-dimensional load data and the baseline response time of each server; for each server, predicting load prediction data and response time prediction data in a next time period according to the multi-dimensional load data of the server; determining a comprehensive load evaluation value of the server according to the load prediction data, the response time prediction data and the index weight coefficients; when the next time period starts, distributing a received service request to a corresponding server in the server cluster for processing according to the comprehensive load evaluation value of each server.
[0008] In a possible implementation, the determining of the index weight coefficient of each load index in the multi-dimensional load data of each server according to the multi-dimensional load data and the benchmark response time comprises: aggregating the multi-dimensional load data of each server to construct a feature matrix, and aggregating the benchmark response time of each server to construct a label vector; performing multivariate linear regression fitting processing on the feature matrix and the label vector by using a least square method to obtain initial weight coefficients of each load index; and performing normalization processing on the initial weight coefficients to obtain the index weight coefficient.
[0009] In a possible implementation, the predicting of the load prediction data and the response time prediction data in the next time period according to the multi-dimensional load data of the server comprises: inputting the multi-dimensional load data of the server into a one-dimensional convolutional neural network layer of a time series prediction model, extracting local correlation features of different load indexes in the time dimension in the multi-dimensional load data through the one-dimensional convolutional neural network layer to obtain a first feature sequence; inputting the first feature sequence into a bidirectional long short-term memory network layer of the time series prediction model, learning long-term forward and reverse dependency relationships of the first feature sequence in time through the bidirectional long short-term memory network layer to obtain a second feature sequence, the second feature sequence comprising feature states of each time step; inputting the second feature sequence into an attention mechanism layer of the time series prediction model, determining attention weights of the feature states of each time step through the attention mechanism layer; performing weighted summation processing on the feature states of each time step according to the attention weights to generate a context vector; and inputting the context vector into an output layer of the time series prediction model, outputting the load prediction data and the response time prediction data by the output layer.
[0010] In a possible implementation, the load prediction data comprises a plurality of load prediction values in a preset time window, and the response time prediction data comprises a plurality of response time prediction values; and the determining of the comprehensive load evaluation value of the server according to the load prediction data, the response time prediction data and the index weight coefficient comprises: determining a load average value of the plurality of load prediction values; performing weighted summation processing on the plurality of load average values according to corresponding weight coefficients to obtain target load data; determining a time average value of the plurality of response time prediction values; and fusing the target load data and the time average value according to a preset proportion to obtain the comprehensive load evaluation value.
[0011] In a possible implementation, the method further includes: periodically collecting an actual benchmark response time of the server; determining a deviation value between the actual benchmark response time and the response time prediction data; and in a case where the deviation value exceeds a preset threshold, updating the index weight coefficient according to the deviation value, so as to re-perform the step of determining the comprehensive load evaluation value of the server according to the load prediction data, the response time prediction data, and the index weight coefficient, by using the updated index weight coefficient.
[0012] In a possible implementation, the method further includes: in a case where the deviation value exceeds the preset threshold in a plurality of continuous time periods, obtaining incremental operation data of each server, the incremental operation data including multi-dimensional load data and a benchmark response time of each server in a latest time period; and performing incremental training on the time series prediction model by using the incremental operation data, to obtain an updated time series prediction model, so as to perform the step of predicting the load prediction data and the response time prediction data in a next time period according to the multi-dimensional load data of the server, by using the updated time series prediction model.
[0013] In a possible implementation, the step of distributing the received service request to a corresponding server in the server cluster for processing according to the comprehensive load evaluation value of each server includes: converting the comprehensive load evaluation value of each server into a request allocation weight, where the request allocation weight is inversely proportional to the comprehensive load evaluation value; determining a request allocation proportion of each server according to the request allocation weight of all servers; and distributing the received service request to a corresponding server in the server cluster for processing according to the request allocation proportion of each server.
[0014] In a second aspect, the present application provides a load balancing device, including:
[0015] The collection module is configured to collect multi-dimensional load data and a benchmark response time of each server in a historical time period in a server cluster; the first determination module is configured to determine an index weight coefficient of each load index in the multi-dimensional load data according to the multi-dimensional load data and the benchmark response time of each server; the prediction module is configured to, for each server, predict load prediction data and response time prediction data in a next time period according to the multi-dimensional load data of the server; the second determination module is configured to determine a comprehensive load evaluation value of the server according to the load prediction data, the response time prediction data, and the index weight coefficient; and the distribution module is configured to, at the start of the next time period, distribute a received service request to a corresponding server in the server cluster for processing according to the comprehensive load evaluation value of each server.
[0016] In a possible implementation, the first determining module is specifically configured to: aggregate the multi-dimensional load data of each server to construct a feature matrix, and aggregate the benchmark response time of each server to construct a label vector; perform multiple linear regression fitting processing on the feature matrix and the label vector by using a least square method to obtain initial weight coefficients of each load index; and perform normalization processing on the initial weight coefficients to obtain the index weight coefficients.
[0017] In a possible implementation, the prediction module is specifically configured to: input the multi-dimensional load data of the server into a one-dimensional convolutional neural network layer of a time series prediction model, extract local correlation features of different load indexes in the time dimension in the multi-dimensional load data through the one-dimensional convolutional neural network layer to obtain a first feature sequence; input the first feature sequence into a bidirectional long short-term memory network layer of the time series prediction model, learn long-term forward and reverse dependency relationships of the first feature sequence in time through the bidirectional long short-term memory network layer to obtain a second feature sequence, the second feature sequence comprising feature states of each time step; input the second feature sequence into an attention mechanism layer of the time series prediction model, determine attention weights of the feature states of each time step through the attention mechanism layer; perform weighted summation processing on the feature states of each time step according to the attention weights to generate a context vector; and input the context vector into an output layer of the time series prediction model, and output the load prediction data and the response time prediction data by the output layer.
[0018] In a possible implementation, the load prediction data comprises a plurality of load prediction values in a preset time window, and the response time prediction data comprises a plurality of response time prediction values; and the second determining module is specifically configured to: determine load average values of the plurality of load prediction values; perform weighted summation processing on the plurality of load average values according to corresponding weight coefficients to obtain target load data; determine time average values of the plurality of response time prediction values; and fuse the target load data and the time average values according to a preset proportion to obtain the comprehensive load evaluation value.
[0019] In a possible implementation, the apparatus further comprises an updating module configured to: periodically collect an actual benchmark response time of the server; determine a deviation value between the actual benchmark response time and the response time prediction data; and in a case where the deviation value exceeds a preset threshold, update the index weight coefficients according to the deviation value, so as to re-perform the step of determining the comprehensive load evaluation value of the server according to the load prediction data, the response time prediction data and the index weight coefficients by using the updated index weight coefficients.
[0020] In a possible implementation, the apparatus further includes a training module configured to: in a case where the bias value exceeds the preset threshold in a plurality of continuous time periods, acquire incremental running data of each server, the incremental running data including multi-dimensional load data and a baseline response time of each server in a latest time period; and perform incremental training on the time series prediction model using the incremental running data to obtain an updated time series prediction model, so as to perform the step of predicting the load prediction data and the response time prediction data in the next time period according to the multi-dimensional load data of the server by using the updated time series prediction model.
[0021] In a possible implementation, the distribution module is specifically configured to: convert the comprehensive load evaluation value of each server into a request distribution weight, where the request distribution weight is inversely proportional to the comprehensive load evaluation value; determine a request distribution proportion of each server according to the request distribution weights of all servers; and distribute the received service request to the corresponding server in the server cluster for processing according to the request distribution proportion of each server.
[0022] In a third aspect, the present application provides a device, comprising: a processor and a memory, the processor is configured to execute a load balancing program stored in the memory to implement the load balancing method of any one of the first aspect.
[0023] In a fourth aspect, the present application provides a storage medium, the storage medium stores one or more programs, the one or more programs can be executed by one or more processors to implement the load balancing method of any one of the first aspect.
[0024] Compared with the prior art, the above technical solutions provided by the embodiments of the present application have the following advantages: the method provided by the embodiments of the present application first provides a comprehensive and quantitative data basis for analysis by collecting historical multi-dimensional load data and baseline response time; then, the dynamic weight coefficients of each load index are determined based on these historical data, so that the weight distribution is more objective and representative; subsequently, the load and response time of each server in the next time period are predicted, so that the decision basis is changed from the "current instantaneous state" to the "future expected state"; finally, the comprehensive load evaluation value is calculated by fusing the prediction data and the weight coefficient, and the request distribution is performed at the beginning of the next time period. Through the scheme, the future bearing pressure of the server can be evaluated in advance, so that the precise scheduling of traffic can be completed before the burst traffic arrives, and the technical problems of the prior art that cannot cope with burst traffic and realize precise optimization due to lack of foresight are fundamentally solved. BRIEF DESCRIPTION OF DRAWINGS
[0025] The accompanying drawings, which are incorporated herein and constitute a part of the specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0026] In order to more clearly illustrate the technical solutions of the embodiments of the present application or the prior art, the accompanying drawings required to be used in the embodiments or prior art description will be briefly introduced as follows. Obviously, those skilled in the art can obtain other drawings from these drawings without any creative effort.
[0027] One or more embodiments are illustrated by the drawings in the accompanying drawings, which do not constitute a limitation on the embodiments. Elements with the same reference numerals in the drawings represent similar elements unless otherwise specified. The drawings in the accompanying drawings do not constitute a proportional limitation.
[0028] Figure 1 An embodiment flow chart of a load balancing method provided by the embodiments of the present application is shown in the accompanying drawings.
[0029] Figure 2 An embodiment flow chart of another load balancing method provided by the embodiments of the present application is shown in the accompanying drawings.
[0030] Figure 3 An embodiment flow chart of another load balancing method provided by the embodiments of the present application is shown in the accompanying drawings.
[0031] Figure 4 An overall flow schematic diagram of a load balancing method provided by the embodiments of the present application is shown in the accompanying drawings.
[0032] Figure 5 An embodiment block diagram of a load balancing device provided by the embodiments of the present application is shown in the accompanying drawings.
[0033] Figure 6 A structural schematic diagram of an electronic device provided by the embodiments of the present application is shown in the accompanying drawings. DETAILED DESCRIPTION
[0034] In order to make the purposes, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without any creative effort fall within the scope of protection of the present application.
[0035] The disclosure below provides many different embodiments or examples for implementing different structures of the application. For the purpose of simplicity, the description below of the specific examples refers only to the features and concepts of each example in connection with the drawings. Of course, they merely illustrate the application and do not in any way limit the scope of the application. Furthermore, the application can be practiced by using a variety of computer-implemented methodologies, mechanisms, articles of manufacture, and / or any suitable materials, as will be apparent to one of ordinary skill in the art. Accordingly, the disclosure below does not represent a breadth of the application, but merely represents representative examples. Additionally, the disclosure below can refer to a number of different examples using a variety of nomenclature. Such nomenclature is used merely for convenience and, in no way, defines the scope of the application.
[0036] To solve the technical problem that the decision mechanism in the prior art lacks foresight, resulting in the system being unable to cope with burst traffic and achieve precise scheduling optimization, the application provides a load balancing method, which can pre-evaluate the future bearing pressure of a server, so that precise scheduling of traffic can be completed before burst traffic arrives.
[0037] Figure 1 An embodiment flowchart of a load balancing method provided by an embodiment of the application is shown in FIG. 1. As shown in FIG. 1, the method comprises the following steps: Figure 1
[0038] Step 101: Collect multi-dimensional load data and benchmark response time of each server in a server cluster in a historical time period.
[0039] Server cluster: refers to a collection of multiple (N) servers (or nodes) connected through a network, which work cooperatively to provide computing services to the outside.
[0040] Multi-dimensional load data: refers to performance indicator data of multiple dimensions collected from a single server node, which is arranged in time sequence to form time series data. Specifically, it can include but is not limited to CPU (Central Processing Unit) usage rate (%), memory usage rate (%), disk I / O (Input / Output) throughput (MB / s or IOPS (Input / Output Operations Per Second)), network bandwidth usage rate (%), etc.
[0041] Benchmark response time: refers to the request response time (usually in milliseconds ms) measured by a benchmark test interface deployed on each server to perform a standardized task. The interface is configured to perform a fixed mixed task (for example, "perform 1000 row database scan operation and perform 1000 floating point operations"), and its response time can comprehensively reflect the actual processing capacity of the server in terms of CPU, memory, disk I / O, etc.
[0042] Historical time period: refers to a specific length of time before the current time, for example, the last 1 hour.
[0043] The embodiment of the present application is implemented by deploying a monitoring agent on each server node. The agent periodically collects the above-mentioned multidimensional load data items at a frequency of seconds or minutes (for example, every 10 seconds), and calls the local benchmark interface to obtain its response time. All data are timestamped and server-identified, and transmitted to the central processing unit for storage, thereby building a historical time series database for subsequent analysis. In order to achieve data quality, the collected raw data can be preprocessed, for example, using the sliding window 3σ (3-Sigma) rule to remove outliers and perform interpolation repair, and normalizing all indicator data and mapping them to the interval [0,1] to eliminate dimensional effects.
[0044] Step 102: Determine the indicator weight coefficient of each load indicator in the multi-dimensional load data according to the multi-dimensional load data of each server and the benchmark response time.
[0045] Indicator Weight Coefficient: A quantitative parameter that indicates the impact and importance of a corresponding load indicator on the server's overall load or overall processing capacity. A larger coefficient indicates a greater impact of changes in that indicator on the overall load.
[0046] In an embodiment of the present application, step 102 may include the following steps: aggregating the multidimensional load data of each server into a feature matrix, and aggregating the benchmark response time of each server into a label vector; performing multivariate linear regression fitting processing on the feature matrix and the label vector using the least squares method to obtain the initial weight coefficient of each load indicator; and normalizing the initial weight coefficient to obtain the indicator weight coefficient.
[0047] Feature matrix: It is a two-dimensional mathematical matrix, in which each row represents a training sample (i.e., the state of a server at a specific historical moment) and each column represents a feature (i.e., a load indicator, such as CPU usage, memory usage, etc.). This matrix systematically organizes scattered multi-dimensional load data as the input of the machine learning model. Label vector: It is a one-dimensional mathematical vector, in which each element corresponds to the target output value of the corresponding row in the feature matrix. In this scheme, each element is the benchmark response time corresponding to its training sample (the state of the server at a certain moment). It serves as the target of model learning and is used to quantify the comprehensive processing capability of the server in this state. Initial weight coefficient: The parameter value obtained by directly solving the least squares method, which reflects the original degree of influence of each load indicator on the benchmark response time.
[0048] In the scheme, first, data preparation is performed, pre-processed data of all servers in a historical time period is extracted, multi-dimensional load data of each server is aggregated according to sample points to construct a feature matrix X, and corresponding benchmark response time values are aggregated to construct a label vector Y. Subsequently, the least square method is used to perform multiple linear regression fitting on the feature matrix X and the label vector Y to obtain a set of initial weight coefficients (θ1, θ2, …, θn) and a bias term θ0. These weight coefficients quantify the original influence degree of each load index on the benchmark response time (i.e. the comprehensive load of the server). Finally, in order to eliminate the dimensional influence and make the weight more comparable, the initial weight coefficients are normalized, for example, using L1 normalization to make the sum of their absolute values equal to 1, and finally the standardized index weight coefficients used for subsequent calculation are obtained.
[0049] The scheme can completely avoid the subjectivity and one-sidedness of the traditional scheme which relies on artificial experience to preset weights, realize the objectivity and quantification of weight coefficient calculation, further through normalization processing, the weight coefficients can truly reflect the relative influence degree of each performance index on the server load under different business scenarios, and provide a scientific basis for subsequent accurate and adaptive load evaluation and scheduling decision.
[0050] Step 103, for each server, according to the multi-dimensional load data of the server, predicting load prediction data and response time prediction data in the next time period.
[0051] Next time period: refers to a future time window immediately after the current time, the length of which is pre-set, for example, 30 seconds in the future.
[0052] Load prediction data: refers to the estimated value of each load index of the server in the future time window obtained by the prediction model.
[0053] Response time prediction data: refers to the estimated value of the benchmark response time of the server in the future time window obtained by the prediction model.
[0054] The embodiment of the application is implemented by a pre-trained CNN-Bi-LSTM-Attention hybrid deep learning model. Among them, the CNN (Convolutional Neural Network) layer is used to extract the local features and short-term patterns of the input data in the time dimension; the Bi-LSTM (Bidirectional Long Short-Term Memory) layer is used to capture the long-term forward and reverse dependencies and periodicity of the load indicators; the Attention (Attention Mechanism) layer is used to assign weights to different time steps and focus on key time points. For each server, the historical multi-dimensional load data (after normalization) of the server in the recent period (such as 60 seconds) is taken as the model input, and the model outputs the predicted value sequence of each load indicator and the baseline response time in the future time period (such as 30 seconds).
[0055] Step 104, according to the load prediction data, the response time prediction data and the index weight coefficient, determining the comprehensive load evaluation value of the server.
[0056] Comprehensive load evaluation value: a single, comprehensive quantitative score, used to represent the expected load level of a server in the upcoming future time period. The lower the value, the lighter the expected load of the server, and the stronger the processing capability.
[0057] In the embodiment of the application, the load prediction data includes a plurality of load prediction values in a preset time window, and the response time prediction data includes a plurality of response time prediction values; step 104 can include the following steps: determining the load average value of the plurality of load prediction values; performing weighted sum processing on the plurality of load average values according to the corresponding weight coefficients to obtain target load data; determining the time average value of the plurality of response time prediction values; fusing the target load data and the time average value according to a preset proportion to obtain the comprehensive load evaluation value.
[0058] Pre-set time window: refers to the length of the predicted future time period in step 103, for example, 30 seconds in the future. This window is a parameter set in advance. Load prediction value: refers to the predicted value of a single load indicator of the server at each specific time point (such as every second) within the future pre-set time window output by the prediction model. Multiple load prediction values constitute the predicted sequence of the indicator within the entire time window. Response time prediction value: refers to the predicted value of the baseline response time of the server at each specific time point within the future pre-set time window output by the prediction model. Multiple response time prediction values constitute the predicted sequence of the baseline response time within the entire time window. Load average value: refers to a single value obtained by performing arithmetic averaging on all prediction values of a certain load indicator within the future pre-set time window, used to represent the overall average load level of the indicator in the future. Target load data: refers to a comprehensive load score obtained by weighting and summing the load average values of each load indicator using the corresponding indicator weight coefficient calculated in step 102. It combines the future average state of multiple load indicators and their relative importance. Time average value: refers to a single value obtained by performing arithmetic averaging on all prediction values of the baseline response time within the future pre-set time window, used to represent the overall average processing capacity of the server in the future. Pre-set ratio: refers to a pre-set fusion coefficient used to balance the contribution of target load data and time average value in comprehensive evaluation.
[0059] In this scheme, for each load indicator, the arithmetic average of all prediction values within the pre-set future time window is calculated to obtain the "load average value" of the indicator; at the same time, the arithmetic average of all prediction values of the baseline response time within the same time window is calculated to obtain the "time average value". Then, using the weight coefficients of each load indicator determined in step 102, the "load average values" of all indicators are weighted and summed to obtain a fused "target load data". Finally, according to a pre-defined "pre-set ratio" (for example, target load data accounts for 70%, and time average value accounts for 30%), the "target load data" and the "time average value" are linearly combined to finally generate a single, quantitative "comprehensive load evaluation value".
[0060] This scheme can effectively smooth out instantaneous fluctuations and more stably reflect the future overall load trend of the server by averaging multiple prediction values within the future time window; by weighting and fusing the multi-dimensional load average values according to the learned weight coefficients, it can ensure that the comprehensive evaluation result matches the actual needs of different business scenarios; finally, by combining the load score with the baseline response time prediction value directly reflecting the processing capacity according to the pre-set ratio, a comprehensive, accurate and reliable future load evaluation value can be obtained, providing a solid data foundation for the load balancer to make optimal scheduling decisions.
[0061] Step 105, at the beginning of the next time period, distribute the received service requests to the corresponding servers in the server cluster according to the comprehensive load evaluation value of each server.
[0062] The embodiment of the present application is the final scheduling execution action. At the end of the current time window and the beginning of the next time window (i.e. the future time period for which the prediction is made), the load balancer obtains the latest comprehensive load evaluation value of all servers in the cluster, and distributes the received service requests to the server cluster according to the evaluation value.
[0063] Specifically, step 105 can include the following steps: converting the comprehensive load evaluation value of each server into a request allocation weight, wherein the request allocation weight is inversely proportional to the comprehensive load evaluation value; determining the request allocation proportion of each server according to the request allocation weight of all servers; and distributing the received service requests to the corresponding servers in the server cluster according to the request allocation proportion of each server.
[0064] Request allocation weight: refers to a positive value calculated for a server, which is used to directly indicate how much proportion of traffic the load balancer should allocate to it. The larger the value, the more requests the server should receive. Inverse proportion: refers to the product of the request allocation weight of a server and its comprehensive load evaluation value remaining a constant (or an approximate constant), i.e. the higher the evaluation value (the heavier the load) of a server, the lower the weight value allocated to it. Request allocation proportion: refers to the share of the request allocation weight of each server in the total weight of all servers in the cluster, usually represented by a percentage value between 0 and 1, and the sum of the proportions of all servers is 1. The proportion directly determines the percentage of the request flow to each server in the total flow.
[0065] The scheme first performs weight conversion: traversing each server in the server cluster, according to the comprehensive load evaluation value calculated by it, the corresponding request allocation weight is calculated through the inverse function, and it is ensured that the server with lower evaluation value (expected load) obtains higher weight. Then, the request allocation weights of all servers are summarized, and the percentage of the weight of each server in the total weight is calculated, that is, the specific request allocation proportion of each server is obtained. Finally, according to this proportion, the load balancer distributes the newly arrived service requests to the corresponding servers in the server cluster for processing by using weighted round robin algorithm and the like. The scheme can convert the comprehensive evaluation value reflecting the future expected load of the server into specific allocation weight and proportion in inverse proportion, so as to directly convert the forward-looking prediction data into executable scheduling instructions, ensure that the server with the lightest load is preferentially connected and receives more traffic, and through the calculation of accurate allocation proportion, the fine-grained on-demand distribution of traffic can be realized, so as to maximize the utilization of cluster resources, avoid any node from being overloaded too early, and finally significantly improve the overall throughput, resource utilization and the ability to cope with sudden traffic of the system.
[0066] The technical scheme provided by the embodiments of the present application first provides comprehensive and quantitative data basis for analysis by collecting historical multi-dimensional load data and benchmark response time; then, the dynamic weight coefficients of each load index are determined based on these historical data, so that the weight distribution is more objective and representative; subsequently, the load and response time of the next time period of each server are predicted, so that the decision basis is changed from "current instantaneous state" to "future expected state"; finally, the comprehensive load evaluation value is calculated by fusing the prediction data and the weight coefficient, and the request distribution is performed at the beginning of the next time period. Through the scheme, the future bearing pressure of the server can be evaluated in advance, so that the accurate scheduling of traffic can be completed before the arrival of sudden traffic, and the technical problems of the prior art that cannot cope with sudden traffic and realize accurate optimization due to lack of foresight are fundamentally solved.
[0067] Figure 2 An embodiment flowchart of another load balancing method provided by the embodiments of the present application is provided. Figure 2 The flowchart shown in Figure 1 On the basis of the flowchart shown, the following steps are included:
[0068] Step 201, inputting the multi-dimensional load data of the server into a one-dimensional convolutional neural network layer of a time series prediction model, extracting the local correlation features of different load indexes in the time dimension in the multi-dimensional load data through the one-dimensional convolutional neural network layer, and obtaining a first feature sequence;
[0069] Step 202: Input the first feature sequence into the bidirectional long short-term memory network layer of the time series prediction model, and learn the long-term positive and negative temporal dependencies of the first feature sequence through the bidirectional long short-term memory network layer to obtain a second feature sequence, where the second feature sequence includes feature states at each time step;
[0070] Step 203: input the second feature sequence into the attention mechanism layer of the time series prediction model, and determine the attention weight of the feature state of each time step through the attention mechanism layer;
[0071] Step 204: Perform weighted summation on the feature states at each time step according to the attention weight to generate a context vector;
[0072] Step 205: Input the context vector to the output layer of the time series prediction model, and the output layer outputs the load prediction data and the response time prediction data.
[0073] For ease of understanding, steps 201-205 are described below in a unified manner:
[0074] One-dimensional convolutional neural network layer: It is a convolutional neural network specially designed for processing sequence data. Its convolution kernel slides only along the time dimension to extract local patterns and correlation features of different features in the input sequence within a short time window.
[0075] Local correlation characteristics: refers to the short-term dependence relationship between multiple load indicators such as coordinated changes, increases and decreases, or resonance between adjacent moments in a short period of time.
[0076] The first feature sequence refers to the output sequence after processing by the one-dimensional convolution layer, which retains the time order of the original data, but the features of each time point have been converted into a higher-dimensional feature representation that can better represent the local pattern.
[0077] Bidirectional long short-term memory network layer (Bi-LSTM): Consists of forward LSTM and backward LSTM, which can simultaneously learn sequence dependencies from both the past and future directions, thereby more comprehensively capturing long-term patterns in time series.
[0078] Long-term positive and negative dependencies refer to the mutual influence between distant time points in a time series. The positive LSTM learns the impact of history on the future (such as the impact of morning peak load on the current state), while the negative LSTM learns the dependency of future states on the past (such as the implications of current state for future trends).
[0079] The second feature sequence refers to the output sequence after processing by the Bi-LSTM layer. The feature state of each time step contains rich information learned from the context of the entire sequence.
[0080] Attention: A network layer that computes weight assignments by calculating the relevance of the query vector to each time-step feature state, automatically assessing and assigning importance weights to different time-steps for the current prediction task.
[0081] Attention weights: Weight values between 0 and 1 computed by the attention mechanism, representing the importance of the corresponding time-step feature state for making accurate predictions, with the sum of all weights being 1.
[0082] Context vector: A fixed-length vector obtained by weighted summing the feature states of all time-steps in the second feature sequence using attention weights. This vector condenses the most critical information in the entire input sequence and is used for the final prediction.
[0083] Output layer: Usually a fully connected neural network layer that maps the context vector condensed with sequence information to the final desired prediction output dimension, i.e., the load prediction data and response time prediction data in the future time window.
[0084] In the embodiments of the present application, the historical multi-dimensional load data of the server is first input into a one-dimensional convolutional neural network (CNN) layer. This layer effectively extracts short-term local correlation features and fluctuation patterns between different load indicators through its convolution kernel operation in the time dimension, and outputs a first feature sequence with enhanced local feature representation. Subsequently, the sequence is input into a bidirectional long short-term memory network layer (Bi-LSTM). This layer fully learns the long-term bidirectional time-dependent relationships (such as periodicity and trend) contained in the sequence through its forward and backward recurrent processing, and outputs a second feature sequence containing rich temporal context information. Then, the attention mechanism layer processes the sequence, automatically calculates the importance of each time-step feature state for predicting future load (attention weights), and generates a context vector that concentrates key information by weighted summing the features of all time steps according to these weights. Finally, the context vector is sent to the output layer for conversion to generate the final future load prediction data and response time prediction data.
[0085] This scheme can fully utilize the advantages of CNN, Bi-LSTM, and Attention through their collaborative work: CNN effectively captures short-term local patterns, Bi-LSTM deeply understands long-term complex dependencies, and Attention focuses on key time point information, thereby achieving multi-dimensional and high-precision prediction of future load; thus greatly improving the credibility of the prediction results, providing a solid data foundation for subsequent precise load scheduling based on prediction, and fundamentally solving the problems of scheduling lag and blindness caused by the lack of effective prediction capability in traditional schemes.
[0086] Figure 3 An embodiment flowchart of another load balancing method provided by embodiments of the present application. Figure 3 The flowchart shown in Figure 1 Based on the flowchart shown, the following steps are included:
[0087] Step 301, periodically collect the actual benchmark response time of the server;
[0088] Step 302, determine the deviation value between the actual benchmark response time and the response time prediction data;
[0089] Step 303, if the deviation value exceeds a preset threshold, update the index weight coefficient according to the deviation value, and re-execute the step of determining the comprehensive load evaluation value of the server according to the load prediction data, the response time prediction data and the index weight coefficient, using the updated index weight coefficient.
[0090] For ease of understanding, the following describes steps 301-303:
[0091] Actual benchmark response time: refers to the response time actually measured by calling the standardized benchmark test interface deployed on each server again after the predicted future time period ends. This value is an objective standard for evaluating prediction accuracy and the real status of the server.
[0092] Deviation value: refers to the difference between the actual benchmark response time and the corresponding response time prediction data (usually the time average) in the prediction period. This value can be quantitatively obtained by mathematical methods such as absolute error, relative error or root mean square error.
[0093] Preset threshold: a critical value set in advance to determine whether the prediction deviation is acceptable. When the deviation value exceeds this threshold, it is considered that the prediction result differs too much from the actual situation, and the system adjustment mechanism needs to be triggered.
[0094] Update the index weight coefficient: refers to the process of dynamically adjusting the initial weight coefficient obtained by linear regression in step 102 according to the calculated prediction deviation. For example, the formula θ_new = θ_old*(1+α*e) can be used for correction, where θ_old is the original weight coefficient, e is the deviation value, α is a learning rate coefficient, and θ_new is the updated weight coefficient.
[0095] In the embodiments of the present application, first, the actual benchmark response time of each server is collected periodically (for example, every 5 minutes) during system operation as a real label for verifying the accuracy of prediction. Then, the deviation value between the actual value and the response time prediction value (such as the time average value) obtained in the previous prediction period is calculated. The system predefines an acceptable deviation threshold (for example, 0.2); when the calculated deviation value exceeds this threshold, it indicates that the current prediction or weight calculation has a large deviation from the actual situation. At this time, the system will dynamically update the weight coefficients of each load index according to the size of the deviation value according to the predetermined rules (such as using a formula containing a learning rate). The updated weight coefficients will be immediately used in the next round of comprehensive load evaluation value calculation, thereby realizing online adaptive adjustment of model parameters.
[0096] The scheme can continuously perceive external environment changes and prediction errors by introducing a closed-loop feedback mechanism of periodic collection, deviation judgment and coefficient update, and dynamically adjust the weight calculation parameters accordingly. This self-correcting ability ensures that the weight coefficients can always accurately reflect the real influence degree of each load index, thereby significantly improving the reliability of the comprehensive load evaluation value and the adaptability and robustness of the system as a whole, realizing the leap from a static model to dynamic and continuous optimization, and effectively guaranteeing the accuracy of the load balancing strategy in long-term operation.
[0097] In another embodiment, the method can further include the steps of: dynamically distributing and loading the updated index weight coefficients to the load balancing gateway, so that the gateway uses the updated index weight coefficients when calculating the comprehensive load evaluation value of each server subsequently.
[0098] Dynamic distribution: refers to the process of actively and real-time pushing the new configuration data (here, the updated index weight coefficients) generated by the data center or management node calculation to the target component (here, the load balancing gateway) through network communication. Loading: refers to the process of loading the new index weight coefficients into the memory and replacing the original old coefficients after the load balancing gateway receives the new index weight coefficients, so that the new coefficients take effect immediately.
[0099] In this embodiment, the step 303 is performed after the updating of the index weight coefficients according to the deviation value is completed. The system actively pushes and deploys the updated index weight coefficient set (i.e., a set of values reflecting the latest importance of each load index such as CPU, memory, I / O, etc.) to the load balancing gateway (such as Nginx, HAProxy, or its extended module) through a predefined communication interface (such as API call) or configuration management channel (such as a configuration center combined with Nacos, etc.). After the gateway receives and loads these new coefficients, in the subsequent running period, the gateway will no longer use the old coefficients when calculating the comprehensive load evaluation value of each server, but will uniformly use the new set of index weight coefficients that have been optimized through feedback.
[0100] This scheme can achieve full-link closed-loop optimization of the core parameters of the load evaluation algorithm by dynamically issuing the updated index weight coefficients to the gateway. This enables the load balancing strategy to adapt to changes in the business model (such as from CPU-intensive to I / O-intensive) from the root, continuously maintaining the scientificity and accuracy of the evaluation criteria, thereby improving the overall adaptive ability and robustness of the system in long-term operation, and achieving higher-level intelligent operation and maintenance.
[0101] In addition, in yet another embodiment, the method can further include the following steps: obtaining incremental running data of each server under the condition that the deviation value exceeds the preset threshold in a plurality of consecutive time periods, the incremental running data including multi-dimensional load data and baseline response time of each server in the most recent time period; incrementally training the time series prediction model using the incremental running data to obtain an updated time series prediction model, so as to use the updated time series prediction model to perform the step of predicting load prediction data and response time prediction data in the next time period according to the multi-dimensional load data of the server.
[0102] The most recent time period refers to a complete historical time window immediately adjacent to the current time, and the time length is a fixed configuration parameter. Incremental running data refers to the running time data in the latest time period generated during the latest running process of the system and not yet used for model training. This data represents the recent real state of the system and has a high degree of timeliness. Incremental training, also known as online learning or continuous learning, does not retrain the model using all historical data, but only fine-tunes the parameters of the existing model using the latest incremental data, so that the model can quickly adapt to new changes in data distribution.
[0103] This embodiment is a further optimization of step 303. When the system monitors that the prediction deviation exceeds the preset threshold in a plurality of consecutive time periods (for example, 3 consecutive periods), it indicates that the prediction model may have failed to accurately capture the current load change pattern. At this time, the system automatically triggers the model updating process: first, the multi-dimensional load data generated by each server in the latest time period (for example, the last 30 minutes) and the corresponding actual benchmark response time are obtained from the storage to form an incremental running data set. Subsequently, the existing CNN-Bi-LSTM-Attention time series prediction model is incrementally trained using this data set as a new training sample to fine-tune its model parameters. After training is completed, the updated time series prediction model is obtained, and the system immediately switches to using the new model for subsequent load and response time prediction.
[0104] This scheme can enable the prediction model to have continuous evolution capability by introducing a model incremental training mechanism based on the continuous deviation criterion. When the system operation mode changes significantly (for example, the business traffic characteristics change, seasonal effects appear), causing the original model to fail, this mechanism can automatically optimize the model parameters using the latest data to quickly adapt to the new environment, thereby significantly improving the accuracy, robustness and adaptability of the prediction model in long-term operation, and fundamentally ensuring the long-term effectiveness of the load balancing strategy based on prediction.
[0105] Figure 4 The overall flowchart of a load balancing method provided by the embodiments of the present application is shown in FIG. 1, which includes the following steps: Figure 4
[0106] Step 1: Multi-dimensional load data acquisition and preprocessing
[0107] 1.1 Data acquisition mechanism
[0108] A distributed monitoring agent is deployed to collect 5 types of core indicators of the server cluster (including N nodes) in real time, and a time series data set is constructed:
[0109] Wherein, i is the server number, t is the collection timestamp (second level), k = 1 represents the CPU usage rate, k = 2 represents the memory usage rate, k = 3 represents the disk I / O throughput, k = 4 represents the network bandwidth usage rate, and k = 5 represents the benchmark interface response time.
[0110] 1.2 Benchmark interface design
[0111] A standardized benchmark test interface is deployed on each server, which fixedly performs a mixed task of "1000 row database scan operation and 1000 times of floating point operation", and uses the response time as a key reference index reflecting the comprehensive processing capability of the server.
[0112] 1.3 Data preprocessing
[0113] Outlier processing: The sliding window 3σ rule is used to detect outliers. Outliers outside the range of μ±3σ are repaired by linear interpolation: (Outlier scenario); Normalization: Use the minimum-maximum normalization method to map each indicator data to the [0,1] interval: in, Indicates the minimum indicator value, Indicates the maximum indicator value.
[0114] Step 2: Dynamic weight calculation based on linear regression
[0115] 2.1 Feature Matrix Construction
[0116] Extract the preprocessed data within the last hour and construct the feature matrix X and label vector Y:
[0117] Input features: (CPU / Memory / I / O / Bandwidth); Target Tags: (Benchmark interface response time).
[0118] 2.2 Linear Regression Modeling
[0119] The influence weight of each indicator on the benchmark response time is solved through the multivariate linear regression model:
[0120] The regression model is expressed as: where w k is the weight coefficient, ∈ is the error term;
[0121] Use the least squares method to solve the optimal weight:
[0122] 2.3 Weight Adaptive Adjustment
[0123] The basic weight coefficients obtained are normalized to achieve standardization of weight coefficients and automatic adaptation to business characteristics:
[0124] When the weight coefficient corresponding to w1′ (CPU usage) increases, the system automatically adapts to CPU-intensive service characteristics; when the weight coefficient corresponding to w3′ or w4′ (I / O / bandwidth weight) increases, the system automatically adapts to I / O-intensive service characteristics.
[0125] Step 3: Load prediction based on CNN-Bi-LSTM-Attention
[0126] 3.1 Model Input and Output
[0127] Model input: historical multidimensional normalized data of each server in the last 60 seconds Model output: predicted values of each load metric and baseline response time within the next 30 seconds
[0128] 3.2 Model structure and calculation
[0129] CNN layer: extract local features through one-dimensional convolution operation: 3) (f is the size of the convolution kernel), used to capture the correlation between adjacent time steps (such as the short-term coupling relationship between CPU and memory); Bi-LSTM layer: learn time series dependency through bidirectional long short-term memory network: (fuse bidirectional features), used to learn long-term time-dependent features of load metrics (such as periodic fluctuation rules); Attention layer: focus on key time step features through attention mechanism: C = ∑a t h t , where a t is the attention weight, C is the weighted context vector, used to enhance the attention on load mutation points (such as request peaks); Output layer: generate prediction values through fully connected layer:
[0130] 3.3 Model training
[0131] Use mean square error loss function: Use Adam optimizer (learning rate set to 0.001) to train the model until the validation set loss is less than 0.01.
[0132] Step 4: comprehensive load evaluation and request scheduling
[0133] 4.1 Prediction value fusion
[0134] Calculate the arithmetic mean of each predicted value within the next 30 seconds as the evaluation basis:
[0135]
[0136] 4.2 Comprehensive score calculation
[0137] Combine dynamic weights and prediction values to calculate comprehensive scores: Where 0.3 is the fixed weight coefficient of the baseline interface response time;
[0138] 4.3 Request allocation weight generation
[0139] Convert the comprehensive score to a scheduling weight: The load balancer distributes requests in proportion to this weight, ensuring that the server with the lowest overall score receives the most requests.
[0140] Step 5: Closed-loop feedback optimization
[0141] 5.1 Deviation Monitoring
[0142] Calculate the deviation between the actual response time and the predicted value every 5 minutes:
[0143] 5.2 Dynamic Correction
[0144] If the deviation value e_i>0.2 (preset threshold), the next round of weight coefficient is corrected according to the formula: w k ′=w k ′×(1-0.1·e i ); If the deviation exceeds the standard for three consecutive times, the prediction model will be incrementally updated using the latest 30-minute data;
[0145] 5.3 Closed-loop feedback implementation
[0146] The updated weights are dynamically written into the load balancing gateway configuration (via Nacos Configuration Center or modules such as nginx_upstream_check_module) to achieve closed-loop feedback and real-time updates of the weight configuration.
[0147] This solution deploys monitoring agents to collect multi-dimensional load and benchmark response time data. After preprocessing, it dynamically calculates weight coefficients for indicators reflecting business characteristics based on linear regression. It then uses a CNN-Bi-LSTM-Attention hybrid model to predict future load trends. It then combines dynamic weights with predicted values to calculate a comprehensive load assessment, which is then used to generate request allocation weights. Finally, through deviation monitoring and dynamic correction mechanisms, it achieves closed-loop optimization of the weight coefficients and prediction model. This solution fundamentally shifts load balancing strategies from passive response to active prediction, effectively improving the system's ability to cope with sudden traffic bursts, resource utilization, and overall stability, while also providing excellent adaptability to business scenarios.
[0148] Figure 5 This is a block diagram of an embodiment of a load balancing device provided in an embodiment of the present application. Figure 5 As shown, the device includes:
[0149] The collection module 51 is used to collect multi-dimensional load data and benchmark response time of each server in the server cluster within a historical time period;
[0150] A first determining module 52 is configured to determine an indicator weight coefficient of each load indicator in the multidimensional load data according to the multidimensional load data of each server and the benchmark response time;
[0151] a prediction module 53, configured to, for each server, predict, according to the multi-dimensional load data of the server, load prediction data and response time prediction data in a next time period;
[0152] a second determination module 54, configured to determine, according to the load prediction data, the response time prediction data and the index weight coefficient, a comprehensive load evaluation value of the server;
[0153] a distribution module 55, configured to, at the start of the next time period, distribute, according to the comprehensive load evaluation value of each server, a received service request to a corresponding server in the server cluster for processing.
[0154] In one possible implementation, the first determination module is specifically configured to: aggregate the multi-dimensional load data of each server to construct a feature matrix, and aggregate the baseline response time of each server to construct a label vector; perform multivariate linear regression fitting processing on the feature matrix and the label vector by using a least square method to obtain initial weight coefficients of each load index; and perform normalization processing on the initial weight coefficients to obtain the index weight coefficient.
[0155] In one possible implementation, the prediction module is specifically configured to: input the multi-dimensional load data of the server into a one-dimensional convolutional neural network layer of a time series prediction model, extract local correlation features of different load indexes in the time dimension in the multi-dimensional load data through the one-dimensional convolutional neural network layer to obtain a first feature sequence; input the first feature sequence into a bidirectional long short-term memory network layer of the time series prediction model, learn long-term forward and reverse dependency relationships of the first feature sequence in time through the bidirectional long short-term memory network layer to obtain a second feature sequence, the second feature sequence containing feature states of each time step; input the second feature sequence into an attention mechanism layer of the time series prediction model, determine attention weights of the feature states of each time step through the attention mechanism layer; perform weighted summation processing on the feature states of each time step according to the attention weights to generate a context vector; and input the context vector into an output layer of the time series prediction model, and output the load prediction data and the response time prediction data by the output layer.
[0156] In a possible implementation, the load prediction data includes a plurality of load prediction values in a preset time window, and the response time prediction data includes a plurality of response time prediction values; the second determining module is specifically configured to: determine load average values of the plurality of load prediction values; perform weighted summation processing on the plurality of load average values according to corresponding weight coefficients to obtain target load data; determine time average values of the plurality of response time prediction values; and fuse the target load data and the time average values according to a preset ratio to obtain the comprehensive load evaluation value.
[0157] In a possible implementation, the apparatus further includes an updating module configured to: periodically collect an actual benchmark response time of the server; determine a deviation value between the actual benchmark response time and the response time prediction data; and in a case where the deviation value exceeds a preset threshold, update the index weight coefficient according to the deviation value, so as to re-perform the step of determining the comprehensive load evaluation value of the server according to the load prediction data, the response time prediction data, and the updated index weight coefficient.
[0158] In a possible implementation, the apparatus further includes a training module configured to: in a case where the deviation value exceeds the preset threshold in a plurality of continuous time periods, acquire incremental running data of the servers, the incremental running data including multi-dimensional load data and a benchmark response time of the servers in a latest time period; and perform incremental training on a time series prediction model by using the incremental running data to obtain an updated time series prediction model, so as to perform the step of predicting the load prediction data and the response time prediction data in a next time period according to the multi-dimensional load data of the servers by using the updated time series prediction model.
[0159] In a possible implementation, the distribution module is specifically configured to: convert the comprehensive load evaluation value of each server into a request allocation weight, where the request allocation weight is inversely proportional to the comprehensive load evaluation value; determine a request allocation proportion of each server according to the request allocation weights of all the servers; and distribute the received service requests to corresponding servers in the server cluster for processing according to the request allocation proportion of each server.
[0160] As shown in Figure 6 The embodiments of the present application provide a device, which includes a processor 111, a communication interface 112, a memory 113, and a communication bus 114, wherein the processor 111, the communication interface 112, and the memory 113 complete mutual communication through the communication bus 114,
[0161] The memory 113 is configured to store a computer program.
[0162] In an embodiment of the present application, the processor 111, when executing the program stored in the memory 113, implements the load balancing method provided by any one of the foregoing method embodiments, including:
[0163] The multi-dimensional load data and the benchmark response time of each server in the server cluster in a historical time period are collected; the index weight coefficients of each load index in the multi-dimensional load data are determined according to the multi-dimensional load data and the benchmark response time of each server; for each server, the load prediction data and the response time prediction data in the next time period are predicted according to the multi-dimensional load data of the server; the comprehensive load evaluation value of the server is determined according to the load prediction data, the response time prediction data and the index weight coefficients; when the next time period starts, the received service requests are distributed to the corresponding servers in the server cluster for processing according to the comprehensive load evaluation values of the servers.
[0164] The present application also provides a computer readable storage medium having a computer program stored thereon, and the computer program, when executed by a processor, implements the steps of the load balancing method provided by any one of the foregoing method embodiments.
[0165] The device embodiments described above are only schematic, wherein the units described as separate components may or may not be physically separate, and the components displayed as units may or may not be physical units, i.e., may be located in one place, or may be distributed on a plurality of network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the present embodiment.
[0166] From the above description of the embodiments, those skilled in the art can clearly understand that the embodiments can be implemented by means of software plus a general hardware platform, and of course can also be implemented by hardware. Based on such understanding, the above technical solutions, essentially or in terms of related art, can be embodied in the form of a software product, and the computer software product can be stored in a computer readable storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, etc., and includes a plurality of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute the methods described in each embodiment or some parts of the embodiments.
[0167] It is to be understood that the terminology used herein is for the purpose of describing particular example embodiments only and is not intended to be limiting. As used herein, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. The terms "comprises", "comprising", "includes", "including" and "has" are inclusive and therefore specify the presence of stated features, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, steps, operations, elements, components, and / or groups thereof. The method steps, processes, and operations described herein are not to be construed as necessarily requiring their performance in the particular order
[0168] The above description is that of current embodiments of the application. Various modifications and changes can be made thereto without departing from the spirit and scope of the application as set forth. The scope of the application is not to be limited to the exact details shown above.
Claims
1. A load balancing method, characterized in that: The method comprises: Collect multi-dimensional load data and benchmark response time of each server in the server cluster within a historical time period; Determining an indicator weight coefficient for each load indicator in the multidimensional load data according to the multidimensional load data of each server and the benchmark response time; For each server, predicting load prediction data and response time prediction data for the next time period based on the multi-dimensional load data of the server; Determining a comprehensive load evaluation value of the server based on the load prediction data, the response time prediction data, and the indicator weight coefficient; At the beginning of the next time period, the received service requests are distributed to the corresponding servers in the server cluster for processing according to the comprehensive load evaluation value of each server.
2. The method according to claim 1, characterized in that Determining the indicator weight coefficient of each load indicator in the multidimensional load data according to the multidimensional load data of each server and the benchmark response time includes: Aggregate the multi-dimensional load data of each server into a feature matrix, and aggregate the benchmark response time of each server into a label vector. Performing a multivariate linear regression fitting process on the feature matrix and the label vector using the least squares method to obtain the initial weight coefficient of each load indicator; The initial weight coefficient is normalized to obtain the indicator weight coefficient.
3. The method according to claim 1, characterized in that The step of predicting load prediction data and response time prediction data within a next time period based on the multi-dimensional load data of the server includes: Inputting the multidimensional load data of the server into a one-dimensional convolutional neural network layer of a time series prediction model, extracting local correlation features of different load indicators in the multidimensional load data in the time dimension through the one-dimensional convolutional neural network layer to obtain a first feature sequence; Inputting the first feature sequence into the bidirectional long short-term memory network layer of the time series prediction model, and learning the long-term positive and reverse dependencies of the first feature sequence in time through the bidirectional long short-term memory network layer to obtain a second feature sequence, wherein the second feature sequence includes feature states of each time step; Inputting the second feature sequence into the attention mechanism layer of the time series prediction model, and determining the attention weight of the feature state of each time step through the attention mechanism layer; Performing weighted summation on the feature states at each time step according to the attention weights to generate a context vector; The context vector is input to the output layer of the time series prediction model, and the output layer outputs the load prediction data and the response time prediction data.
4. The method according to claim 1, wherein The load prediction data includes a plurality of load prediction values within a preset time window, and the response time prediction data includes a plurality of response time prediction values; Determining the comprehensive load evaluation value of the server according to the load prediction data, the response time prediction data, and the indicator weight coefficient includes: Determining a load average of a plurality of the load prediction values; Performing weighted summation processing on the plurality of load average values according to corresponding weight coefficients to obtain target load data; determining a time average of a plurality of the response time prediction values; The target load data and the time average value are fused according to a preset ratio to obtain the comprehensive load evaluation value.
5. The method according to claim 1, wherein The method further comprises: Periodically collecting the actual benchmark response time of the server; Determining a deviation value between the actual benchmark response time and the response time prediction data; When the deviation value exceeds a preset threshold, the indicator weight coefficient is updated according to the deviation value, so as to re-execute the step of determining the comprehensive load evaluation value of the server based on the load prediction data, the response time prediction data and the indicator weight coefficient using the updated indicator weight coefficient.
6. The method according to claim 5, characterized in that The method further comprises: When the deviation value exceeds a preset threshold value in multiple consecutive time periods, obtaining incremental operation data of each server, wherein the incremental operation data includes multi-dimensional load data and benchmark response time of each server in the most recent time period; The incremental running data is used to incrementally train the time series prediction model to obtain an updated time series prediction model, and the updated time series prediction model is used to execute the step of predicting the load prediction data and response time prediction data in the next time period based on the multi-dimensional load data of the server.
7. The method according to claim 1, characterized in that The step of distributing the received service request to the corresponding server in the server cluster for processing according to the comprehensive load evaluation value of each server includes: Converting the comprehensive load evaluation value of each server into a request allocation weight, wherein the request allocation weight is inversely proportional to the comprehensive load evaluation value; Determine the request allocation ratio of each server based on the request allocation weights of all servers; According to the request allocation ratio of each server, the received business request is distributed to the corresponding server in the server cluster for processing.
8. A load balancing device, characterized in that: The device comprises: The collection module is used to collect multi-dimensional load data and benchmark response time of each server in the server cluster within a historical period; A first determining module is configured to determine an indicator weight coefficient of each load indicator in the multidimensional load data according to the multidimensional load data of each server and a benchmark response time; A prediction module, configured to predict, for each server, load prediction data and response time prediction data for a next time period based on the multi-dimensional load data of the server; A second determining module is configured to determine a comprehensive load evaluation value of the server based on the load prediction data, the response time prediction data, and the indicator weight coefficient; The distribution module is used to distribute the received business requests to the corresponding servers in the server cluster for processing according to the comprehensive load evaluation value of each server at the beginning of the next time period.
9. An electronic device, characterized in that: include: A processor and a memory, wherein the processor is configured to execute a load balancing program stored in the memory to implement the load balancing method according to any one of claims 1 to 7.
10. A storage medium, characterized in that: The storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the load balancing method according to any one of claims 1 to 7.
Citation Information
Cited By
Server cluster load balancing state tracking analysis method
CN121711349A