Deep Learning-Based Server Load Balancing Prediction Method and System
By combining topological embedding, group theory transformation, spectral decomposition and deep learning technologies, capturing the geometry and symmetry of server load data, the existing methods are solved in terms of prediction accuracy, efficiency and interpretability, and efficient load balancing prediction and resource optimization are achieved.
Patent Information
- Application Number
- CN202510048295.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-13
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-01-13
AI Technical Summary
When handling complex and variable load modes, existing server load prediction methods are difficult to capture long-term dependencies and nonlinear features, insufficient prediction accuracy, lack consideration of multidimensional characteristics and interactions, low computational efficiency, and lack interpretability.
Using a deep learning-based method, combining topological embedding, group theory transformation, spectral decomposition and long-term memory networks, the geometric structure and symmetry of load data are captured through Riemann manifold mapping, eigenvalue decomposition and nonlinear transformation, multi-scale time dependencies are modeled, and resource allocation schemes are generated through multi-objective optimization algorithms.
It improves prediction accuracy and computing efficiency, enhances the interpretability and adaptability of the model, and can accurately capture the complex patterns and long-term trends of server load, providing a reliable basis for load balancing strategies.
Smart Images

Figure CN119484541B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of server load prediction, and more specifically, to a server load balancing prediction method and system based on deep learning. Background Art
[0002] With the rapid development of cloud computing and big data technologies, the scale of data centers has been continuously expanding, and server load balancing prediction has become a key technology to ensure the stable operation of the system and improve resource utilization efficiency. Traditional load balancing prediction methods mainly rely on statistical models and simple machine learning algorithms, such as the moving average method, exponential smoothing method, and autoregressive model, etc. These methods perform okay when dealing with linear and short-term trends, but when facing the complex and changeable load patterns in modern data centers, they often have difficulty capturing long-term dependencies and non-linear characteristics, resulting in insufficient prediction accuracy.
[0003] In recent years, with the rise of deep learning technologies, some researchers have attempted to apply recurrent neural networks (RNNs) and long short-term memory networks (LSTMs) to server load prediction. These methods have improved the prediction accuracy to a certain extent, especially showing obvious advantages in dealing with time series data. However, these methods still have some significant technical problems: First, they usually directly use the original load data as input, ignoring the inherent geometric structure and symmetry of the data, resulting in the model being difficult to fully utilize the implicit information in the data; Second, these methods have low computational efficiency when dealing with high-dimensional data and are difficult to meet the requirements of real-time prediction in large-scale data centers; Third, existing methods generally lack effective modeling of multi-scale time dependencies and are difficult to capture short-term fluctuations and long-term trends simultaneously.
[0004] In addition, existing load prediction methods often regard the prediction problem as a single regression task, ignoring the multi-dimensional characteristics of load data and the interaction between different features. This results in the prediction results being difficult to comprehensively reflect the actual operating state of the server, affecting the formulation and execution of load balancing strategies. At the same time, most existing methods lack interpretability of the prediction results, making it difficult for system administrators to understand and verify the decision-making basis of the prediction model, reducing the credibility of the prediction results in practical applications.
[0005] In view of the above problems, there is an urgent need for a load balancing prediction method that can comprehensively consider the characteristics of server load data and has high accuracy, high efficiency, and good interpretability. The present invention is an innovative solution proposed in response to this need. Summary of the Invention
[0006] The technical problem to be solved by the present invention is how to improve the prediction efficiency while ensuring the prediction accuracy, enhance the interpretability and adaptability of the model, so as to achieve more intelligent and efficient server load balancing. Specifically, the present invention is committed to solving the following key problems: how to effectively capture the geometric structure and symmetry of the load data; how to reduce the computational complexity while maintaining high-dimensional information; how to model multi-scale time dependencies simultaneously; how to improve the interpretability and credibility of the prediction results.
[0007] The present invention provides a server load balancing prediction method based on deep learning, including:
[0008] An acquisition step, including:
[0009] Obtaining the real-time load data and historical load data of the server;
[0010] A processing step, including:
[0011] Performing a topological embedding transformation based on the real-time load data and the historical load data;
[0012] Performing a group theory transformation according to the result of the topological embedding transformation;
[0013] Performing a spectral decomposition based on the result of the group theory transformation;
[0014] Performing a non-linear transformation according to the result of the spectral decomposition;
[0015] Performing a time series prediction using a long short-term memory network based on the result of the non-linear transformation;
[0016] An output step, including:
[0017] Outputting the future load prediction result of the server.
[0018] Preferably, the topological embedding transformation specifically includes:
[0019] Mapping the server load data onto a Riemannian manifold;
[0020] Using the Riemannian exponential map to achieve the topological embedding of the data.
[0021] Preferably, the group theory transformation specifically includes:
[0022] Selecting the special orthogonal group SO(n) as the transformation group;
[0023] Achieving the rotational invariance of the data through group action.
[0024] Preferably, the spectral decomposition specifically includes:
[0025] Performing eigenvalue decomposition on the data after the group theory transformation;
[0026] Obtain the eigenvector matrix and the eigenvalue diagonal matrix.
[0027] Preferably, the non-linear transformation specifically includes:
[0028] Perform non-linear mapping on the spectral decomposition result using an activation function;
[0029] Calculate the trace of the transformed matrix as the output.
[0030] Preferably, the long short-term memory network performing time series prediction specifically includes:
[0031] Process the data sequence after non-linear transformation using an LSTM network;
[0032] Output the final prediction result through a softmax function.
[0033] Preferably, it further includes a data preprocessing step:
[0034] Perform normalization processing on the obtained load data;
[0035] Remove outliers and noise data.
[0036] Preferably, it further includes a model training step:
[0037] Train the long short-term memory network using historical load data;
[0038] Optimize the network parameters through the backpropagation algorithm.
[0039] Preferably, it further includes a load balancing strategy formulation step:
[0040] Generate a server resource allocation plan based on the future load prediction result;
[0041] Adjust the load distribution of the server cluster according to the resource allocation plan.
[0042] A server load balancing prediction system based on deep learning for executing the method includes:
[0043] A data acquisition module, configured to acquire real-time load data and historical load data of a server;
[0044] A topology embedding module, configured to perform topology embedding transformation based on the real-time load data and the historical load data;
[0045] A group theory transformation module, configured to perform group theory transformation according to the result of the topology embedding transformation;
[0046] A spectral decomposition module, configured to perform spectral decomposition based on the result of the group theory transformation;
[0047] A non - linear transformation module, which is used to perform a non - linear transformation according to the result of spectral decomposition;
[0048] A time - series prediction module, which is used to perform time - series prediction using a long short - term memory network based on the result of the non - linear transformation;
[0049] An output module, which is used to output the prediction result of the future load of the server.
[0050] The beneficial effects of the present invention are mainly reflected in the following aspects:
[0051] The present invention innovatively combines topology, group theory and deep learning technologies to propose a brand - new server load balancing prediction method. This method not only significantly outperforms the prior art in terms of prediction accuracy, but also makes important breakthroughs in terms of computational efficiency, model adaptability and result interpretability.
[0052] First of all, by introducing topological embedding and group - theory transformation, the present invention effectively captures the geometric structure and symmetry of the load data. This pre - processing not only improves the quality of feature representation, but also enhances the invariance of the model to data rotation and translation, thus improving the robustness and generalization ability of the prediction. Especially when dealing with load patterns with periodicity and symmetry, the method of the present invention shows obvious advantages.
[0053] Secondly, the present invention combines spectral decomposition and non - linear transformation. While retaining the high - dimensional information of the data, it effectively reduces the computational complexity. This dimensionality reduction method is different from traditional principal component analysis (PCA). It can retain the non - linear features of the data and provide richer and more meaningful inputs for subsequent deep - learning models. Experimental results show that this processing method not only improves the prediction accuracy, but also significantly reduces the calculation time, enabling the method of the present invention to meet the real - time prediction requirements of large - scale data centers.
[0054] Furthermore, through the multi - layer LSTM network structure, the present invention successfully models the multi - scale time - dependent relationship of the load data. This structure can capture both short - term fluctuations and long - term trends, making the prediction results more comprehensive and accurate. Especially when predicting load changes over a long time span (such as 24 hours or longer), the method of the present invention shows obvious advantages, which is of great significance for formulating long - term resource scheduling strategies.
[0055] In addition, the method of the present invention has good interpretability. By analyzing the results of topological embedding and group theory transformation, system administrators can intuitively understand the structural characteristics of the data; by observing the results of spectral decomposition, key factors affecting load changes can be identified; by analyzing the attention mechanism of the LSTM network, the key time periods that the model focuses on when making predictions can be understood. This interpretability not only enhances the credibility of the prediction results but also provides valuable insights for system optimization and fault diagnosis.
[0056] Finally, the method of the present invention has strong adaptability and scalability. Through the continuous learning mechanism, the model can automatically adapt to the dynamic changes of the load pattern; through the modular design, the system can flexibly adjust or replace certain components according to specific requirements to adapt to data centers of different scales and types.
[0057] In summary, the server load balancing prediction method and system based on deep learning proposed by the present invention have made significant breakthroughs in multiple aspects such as prediction accuracy, computational efficiency, model adaptability, and result interpretability. These innovations not only solve the key problems faced by the existing technologies but also open up new possibilities for the intelligent management of data centers, and are expected to play an important role in improving resource utilization efficiency, reducing energy consumption, and ensuring service quality, providing strong technical support for server management in the cloud computing and big data era. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] Figure 1 is the system flow chart of the present invention.
[0059] Figure 2 is the logical block diagram of the data acquisition module of the present invention.
[0060] Figure 3 is the detailed flow chart of the processing steps of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0061] Please refer to Figures 1 - 3 , the present invention relates to a server load balancing prediction method and system based on deep learning. By innovatively combining topology, group theory, and deep learning technologies, the method realizes high-precision prediction of server loads, providing a reliable basis for the optimal allocation of server resources.
[0062] First, the method of the present invention includes an acquisition step, a processing step, and an output step. In the acquisition step, the method acquires real-time load data and historical load data of the server. These data typically include key performance indicators (KPIs) such as CPU usage rate, memory usage rate, network traffic, etc. Preferably, the present invention adopts a high-frequency sampling method, such as collecting data once per second, to ensure capturing the subtle trends of load changes. This ensures the accuracy of subsequent analysis. By comprehensively analyzing multiple KPIs, a more comprehensive view of the server's operation can be obtained, which helps to discover hidden problem patterns or abnormal situations.
[0063] In the processing step, the method first performs a topological embedding transformation based on the acquired real-time load data and historical load data. This innovative step maps the multi-dimensional load data into a topological space with geometric significance. Specifically, the present invention adopts the following topological embedding function:
[0064] ,
[0065] where is d-dimensional Euclidean space, is the number of server load characteristics, is a Riemannian manifold. The specific implementation of this function is as follows:
[0066] ,
[0067] Here, is the Riemannian exponential map, is a fixed point on the manifold, is an orthonormal basis in the tangent space, is the th eigenvalue of the load data. Through this mapping, while maintaining the data's topological structure, the high-dimensional data can be transformed into a more easily processed low-dimensional representation.
[0068] Here, it is assumed that d-dimensional Euclidean space the vector in represents the set of load characteristics of a server at a certain point in time (such as CPU usage rate, memory usage rate, etc.). The Riemannian manifold is a space with a specific geometric structure. The mapping transforms these characteristics to the manifold through the Riemannian exponential map where is a fixed point on the manifold, and is an orthonormal basis in the tangent space.
[0069] For example, in a data center, if each server has 20 different performance metrics, the dimension of the original data is very high. Through topological embedding transformation, these 20 features can be mapped onto a low-dimensional manifold that preserves the intrinsic structure of the data, thereby reducing the computational complexity. Consider two servers A and B with similar load characteristics at different time points. After topological embedding transformation, their representations on the manifold remain close, ensuring that the model can capture similar working patterns or anomalies.
[0070] Next, the method performs a group theory transformation based on the result of the topological embedding transformation. The purpose of this step is to capture the symmetry in the data and improve the generalization ability of the model. The present invention selects the special orthogonal group SO(n) as the transformation group, where is the dimension of the topological space. The group action is defined as follows:
[0071] ,
[0072] where is a group element that can be learned through the network. This transformation can achieve the rotational invariance of the data, making the model robust to rotational changes in the data.
[0073] For the group element , where is the dimension of the topological space, it acts on the data that has already undergone topological embedding transformation. This operation actually changes the perspective by rotation while keeping the internal relationships of the data unchanged.
[0074] Imagine a periodic load change pattern (such as the difference between weekdays and weekends) existing in a data center. Through group theory transformation, even if the data is rotated (i.e., observed from different angles), the model can still identify the same pattern, improving the stability and reliability of the prediction. For example, when the load distribution changes due to hardware failures of some machines in a server cluster, group theory transformation can help extract the symmetric characteristics behind this change, enabling the model to more effectively adapt to the new load pattern.
[0075] After the group theory transformation, the method performs spectral decomposition based on the transformation result. Spectral decomposition is a powerful technique that can reveal the intrinsic structure of the data.
[0076] After completing the spectral decomposition, this method performs a non - linear transformation according to the decomposition result. This step introduces non - linearity and enhances the model's expressive power. This non - linear transformation can capture complex patterns in the data and improve the model's fitting ability. Finally, based on the result of the non - linear transformation, this method uses a Long Short - Term Memory network (LSTM) to perform time - series prediction. The LSTM network is a special type of recurrent neural network, which is particularly suitable for processing time - series data.
[0077] In the output step, this method outputs the prediction results of the future server load. These results can be the load trends within a future period (such as the next 1 hour, 24 hours, etc.), providing a basis for load - balancing decisions.
[0078] The method of the present invention realizes high - precision prediction of server load by innovatively combining topology, group theory, and deep - learning techniques. Compared with traditional methods, this method has the following advantages: First, the topological embedding transformation preserves the geometric structure of the data, which helps to capture the internal laws of load changes; Second, the group - theory transformation enhances the rotational invariance of the model and improves the stability of prediction; Third, spectral decomposition and non - linear transformation effectively extract the key features of the data; Finally, the LSTM network well handles the time - series dependence of load data. The synergistic effect of these innovation points enables this method to more accurately predict server load and provides strong support for achieving efficient load balancing.
[0079] In a preferred embodiment of the present invention, the spectral decomposition specifically includes performing eigenvalue decomposition on the data after group - theory transformation and obtaining the eigenvector matrix and the eigenvalue diagonal matrix. This step is a further processing of the result of the aforementioned group - theory transformation, aiming to reveal the internal structure and main components of the data.
[0080] Specifically, the spectral - decomposition operation can be expressed as:
[0081] ,
[0082] where, represents the spectral - decomposition operation, is the data after group - theory transformation, is the eigenvector matrix, is the eigenvalue diagonal matrix. Through this decomposition, the method of the present invention can decompose the complex data structure into a series of orthogonal basis vectors (eigenvectors) and their corresponding importance weights (eigenvalues).
[0083] Here, is the eigenvector matrix, is the eigenvalue diagonal matrix. Through spectral decomposition, the main components and change patterns of the data can be obtained, providing important information for subsequent prediction tasks.
[0084] Spectral decomposition is the process of performing eigenvalue decomposition on the data matrix after group theory transformation. Suppose is the data matrix after group theory transformation, then it can be decomposed into , where is an orthogonal matrix composed of eigenvectors, is a diagonal matrix containing the eigenvalues sorted by magnitude.
[0085] In a large data center, there may be a dataset with thousands of records. By selecting the eigenvectors corresponding to the largest several eigenvalues, the dimensionality of the data can be significantly reduced while almost no useful information is lost. For example, retaining only the first 50 principal components may be sufficient to describe more than 95% of the variance. In a practical environment, the server load data will inevitably contain some random fluctuations or measurement errors. Spectral decomposition helps remove the parts corresponding to these small eigenvalues, reducing noise interference and improving prediction accuracy.
[0086] Consider a data center with 100 servers, each with 20 performance metrics, forming a 100×20 data matrix. After spectral decomposition, it is found that the first 10 eigenvalues account for 98% of the total variance. Therefore, the corresponding 10 eigenvectors can be selected as the basis of the new feature space, forming a 100×10 reduced-dimensional data representation. This not only simplifies the problem but also improves the computational efficiency.
[0087] There may be a large amount of redundant information or noise in the load data of the data center. The eigenvector matrix helps isolate the main components, enabling the model to focus on the truly meaningful patterns without being disturbed by irrelevant details.
[0088] For example, the load changes of some servers may be caused by the same type of user activities, such as accessing specific web pages or running similar applications. The eigenvector matrix can identify these common behavior patterns and classify them into a few important eigenvectors, thereby reducing the sensitivity of the model to noisy data.
[0089] The eigenvalue diagonal matrix provides a quantitative way to evaluate the importance of each eigenvector. By selecting the eigenvectors corresponding to the largest several eigenvalues, the dimensionality of the data can be significantly reduced while almost no useful information is lost.
[0090] Continuing with the data center example mentioned earlier, if you want to further compress the data, you can set a threshold (such as 95% of the total variance), and then select all the eigenvalues greater than this threshold and their corresponding eigenvectors. This can not only significantly reduce the amount of data but also ensure that the most critical information is retained.
[0091] Small eigenvalues typically correspond to random fluctuations or measurement errors in the data. By ignoring these eigenvalues and their corresponding eigenvectors, noise can be effectively removed and the prediction accuracy can be improved.
[0092] In actual operation, the data center may encounter problems such as hardware failures and network jitters, resulting in abnormal fluctuations in the load data of some servers. The eigenvalue diagonal matrix helps identify and filter out these atypical data points, enabling the model to more robustly capture the normal working mode.
[0093] On a large cloud computing platform, an administrator hopes to adjust the resource allocation plan of the server cluster through a multi-objective optimization algorithm. In this process, spectral decomposition plays a crucial role:
[0094] Initial stage: Collect the server load data over a period of time and construct a high-dimensional data matrix .
[0095] Spectral decomposition: Perform spectral decomposition on the data matrix to obtain the eigenvector matrix and the eigenvalue diagonal matrix .
[0096] Dimensionality reduction and information concentration: Select the most important principal components according to the eigenvalue magnitudes and construct a new low-dimensional data representation.
[0097] Multi-objective optimization: Use the reduced-dimensional data as input and generate multiple non-dominated solutions (Pareto optimal solutions) through multi-objective optimization algorithms such as NSGA-II, providing multiple feasible options for decision-makers. Through the above steps, not only is the problem complexity simplified and the computational efficiency improved, but also the model's ability to understand the real-world load patterns is enhanced, ultimately achieving more efficient and stable resource management and load balancing.
[0098] Preferably, the present invention uses the singular value decomposition (SVD) algorithm to implement spectral decomposition. The SVD algorithm has the characteristics of good numerical stability and high computational efficiency, and is particularly suitable for processing large-scale server load data. In actual applications, complete SVD or truncated SVD can be selected according to specific requirements to achieve a balance between computational efficiency and information retention.
[0099] When the system of the present invention performs spectral decomposition, it usually retains the principal components that account for more than 95% of the total variance. This threshold is an empirical value obtained through a large number of experiments, which can effectively reduce the data dimensionality while retaining the key information and improving the efficiency of subsequent processing.
[0100] Next, the method of the present invention performs a non - linear transformation based on the result of spectral decomposition. This step introduces non - linearity and significantly enhances the expressive power of the model. The specific form of the non - linear transformation is as follows:
[0101] ,
[0102] In this formula, represents the non - linear transformation function, is the activation function, tr represents the trace operation of the matrix, is the weight matrix, is the bias term. The present invention preferably uses ReLU (Rectified Linear Unit) as the activation function because ReLU has advantages such as simple calculation and stable gradient, and is particularly suitable for the training of deep networks.
[0103] The non - linear transformation introduces the activation function (such as ReLU), the weight matrix , the bias , and the trace operation tr. The trace operation is used to compress a high - dimensional matrix into a scalar output while retaining the important features of the matrix. In this example, is the matrix after spectral decomposition.
[0104] Taking a complex server load scenario as an example, where there are not only regular periodic fluctuations but also sudden traffic peaks. The non - linear transformation allows the model to fit such complex functional relationships and improves the ability to capture such load patterns. By adjusting the weight matrix , the model can dynamically control the importance of each feature according to the knowledge learned during the training process. For example, if it is found that a certain feature (such as network latency) is particularly important for predicting future load trends, the prediction performance can be optimized by increasing the weight corresponding to this feature.
[0105] It is worth noting that the present invention uses the trace operation of the matrix in the non - linear transformation. This is an innovative point, which can effectively compress a high - dimensional matrix into a scalar while retaining the important features of the matrix. By adjusting the weight matrix , the method of the present invention can flexibly control the importance of each feature.
[0106] In an embodiment of the present invention, the initial value of the weight matrix W is set using the Xavier initialization method. This initialization method can make the output variance of each layer as equal as possible, help prevent the problem of gradient vanishing or explosion, and accelerate the convergence of the network.
[0107] After completing the non - linear transformation, the method of the present invention uses a Long Short - Term Memory network (LSTM) to perform time - series prediction. The LSTM network is a special type of recurrent neural network, and its core lies in the introduction of gate mechanisms, which can effectively handle long - term dependence problems. The prediction process can be expressed as:
[0108] ,
[0109] ,
[0110] Here, is the hidden state, is the cell state, is the prediction output. Through the LSTM network, this method can effectively capture the long - term dependence relationship of the load data and improve the accuracy of prediction. The internal structure of the LSTM cell includes an input gate, a forget gate, and an output gate. These gated units enable the LSTM to dynamically adjust the retention and forgetting of information according to the importance of the current input and historical information.
[0111] The Long Short - Term Memory network (LSTM) is a recurrent neural network specifically designed to process sequential data. The LSTM cell contains an input gate, a forget gate, and an output gate inside, which work together to determine which information should be remembered or forgotten. The prediction output is calculated from the hidden state at the current time step through a softmax layer.
[0112] Considering that the load of the data center usually has a certain periodicity and trend (such as a sharp increase in traffic during the morning rush hour every day), the LSTM is particularly good at dealing with dependencies over a long time span. This enables it to accurately predict the load trend in the next few hours or even days.
[0113] Preferably, the LSTM network of the present invention adopts a bidirectional structure. The bidirectional LSTM structure can consider both past and future context information simultaneously, which is particularly effective for capturing the periodic patterns and long - term trends of server loads. For example, it can better identify the load differences between weekdays and weekends, or capture the load changes during special periods such as holidays.
[0114] During the network training process, the present invention adopts the Adam optimizer, with the initial learning rate set to 0.001, and a learning rate decay strategy is used. Specifically, every 10 epochs, the learning rate decays to 0.9 times the original value. This strategy can quickly approach the optimal solution in the initial stage of training and finely adjust the parameters in the later stage of training to improve the performance of the model.
[0115] In addition, to prevent overfitting, the present invention introduces the Dropout regularization technique in the LSTM network. The Dropout rate is set to 0.5, which is a valid value widely verified in practice and can significantly improve the generalization ability of the model while maintaining its expressive power.
[0116] Through the above series of innovative processing steps, the method of the present invention can comprehensively and deeply analyze the characteristics of server load data, capture the complex patterns and long-term dependencies therein, so as to achieve high-precision load prediction. This provides reliable data support for the optimal allocation of server resources and the formulation of load balancing strategies, and helps improve the operation efficiency and service quality of the data center. In another preferred embodiment of the present invention, the method further includes a model training step. This step is crucial for ensuring the performance of the prediction model. Specifically, the present invention uses historical load data to train the long short-term memory network and optimizes the network parameters through the backpropagation algorithm.
[0117] During the model training process, the present invention adopts the batch gradient descent method, and the number of samples in each batch is set to 64. This batch size is an empirical value obtained through multiple experiments and can achieve a good balance between training stability and computational efficiency. The training data is divided into a training set, a validation set, and a test set according to a ratio of 8:1:1. This division method can make full use of the data and ensure that the generalization ability of the model is effectively evaluated.
[0118] Preferably, the present invention introduces an early stopping strategy during the training process. Specifically, if the loss on the validation set does not decrease for 10 consecutive epochs, the training is stopped. This strategy can effectively prevent overfitting and save computational resources. In addition, the present invention also adopts a learning rate adaptive adjustment technique. The initial learning rate is set to 0.001, and when the validation set loss does not decrease for 3 consecutive epochs, the learning rate will automatically be reduced to half of the original. This dynamic adjustment strategy can help the model converge to the optimal solution faster.
[0119] After the training is completed, the present invention evaluates the model performance on the test set. The evaluation metrics include root mean square error (RMSE), mean absolute error (MAE), and mean absolute percentage error (MAPE). Usually, the method of the present invention can control the MAPE within 5%, which means that the prediction results have high practical value.
[0120] An important feature of the present invention is that model training is not a one-time process. In practical applications, the system of the present invention will regularly (e.g., weekly) retrain the model using newly collected data. This continuous learning mechanism enables the model to adapt to changes in server load patterns in a timely manner and maintain the accuracy of predictions.
[0121] After completing the load prediction, the method of the present invention further includes a load balancing strategy formulation step. This step converts the prediction result into an actual resource scheduling decision, which is a key link in the practical application of the present invention.
[0122] Specifically, based on the future load prediction result, the present invention generates a server resource allocation plan. This process involves the balance of multiple factors, including but not limited to: the predicted load level, the current state of the server, energy consumption efficiency, service quality requirements, etc. The present invention uses a multi-objective optimization algorithm to generate the resource allocation plan, and its objective function can be expressed as:
[0123] ,
[0124] where, represents the resource allocation plan, represents the th optimization objective (such as load balancing degree, energy consumption, response time, etc.). The present invention preferably uses the NSGA-II (Non-dominated Sorting Genetic Algorithm II) algorithm to solve this multi-objective optimization problem. The NSGA-II algorithm can obtain multiple non-dominated solutions in one run, providing multiple optional solutions for decision-makers.
[0125] Each optimization objective is a function that measures a specific performance index, and they jointly determine the quality of the final resource allocation plan. These objectives may include but not limited to:
[0126] Load balancing degree: Ensure that the workload on each server is as evenly distributed as possible.
[0127] Energy consumption efficiency: Minimize the power consumption of the entire data center.
[0128] Response time: Minimize the average processing time of user requests as much as possible.
[0129] Quality of Service (QoS): Ensure that the priorities of critical tasks and services are met.
[0130] For the above four objectives, the following specific mathematical expressions can be defined:
[0131] Load balancing degree: Assume is the number of servers, is the load of the th server, then the load balancing degree can be measured by the standard deviation :
[0132] ,
[0133] where, is the average load of all servers.
[0134] Energy consumption efficiency: Let be the power consumption of the th server at time, then the total energy consumption can be given by integration or summation:
[0135] or , Response time: Let be the response time of the th request, then the average response time is:
[0136] ,
[0137] where is the number of requests.
[0138] Quality of service: Considering the importance and urgency of different services, a weighting factor can be introduced to evaluate the overall quality of service
[0139] ,
[0140] where is the quality score of the th service.
[0141] The present invention preferably uses NSGA-II (Non-dominated Sorting Genetic Algorithm II) to solve the above multi-objective optimization problem. NSGA-II is an evolutionary algorithm that can find a set of non-dominated solutions (i.e., Pareto optimal solutions) in one run, thus providing multiple feasible options for decision-makers.
[0142] The working principle of NSGA-II is as follows:
[0143] Initialize the population: Randomly generate a set of initial resource allocation schemes as the first-generation population.
[0144] Evaluate individual fitness: Calculate the values of each scheme for each objective function and classify the individuals according to fast non-dominated sorting.
[0145] Selection operation: Select excellent individuals from the current population to participate in crossover and mutation to generate the next-generation population.
[0146] Crossover and mutation: Perform gene exchange (crossover) and random change (mutation) on the selected individuals to create new solutions.
[0147] Environmental Selection: Combine the parent and offspring populations, and then select the members of the next generation population through crowding distance sorting.
[0148] Termination Condition: When the preset maximum number of iterations is reached or other stopping criteria are met, output the optimal solution set.
[0149] On a large cloud computing platform, an administrator hopes to optimize the load balancing degree, energy consumption efficiency, response time, and service quality simultaneously. By using the NSGA-II algorithm, the system can generate a series of different resource allocation schemes, and each scheme corresponds to a specific set of optimization results. For example:
[0150] Scheme A: Focus on load balancing degree and energy consumption efficiency, suitable for use during daily operations to maintain low costs and stable performance.
[0151] Scheme B: Emphasize response time and service quality, applicable during peak periods or important business hours to ensure that the user experience is not affected.
[0152] Scheme C: Balance various indicators and provide a comprehensively optimized resource configuration under normal circumstances.
[0153] Using a multi-objective optimization algorithm to generate resource allocation schemes has the following significant advantages:
[0154] Flexibility: According to the changing demands in different time periods, the most suitable optimization strategy for the current situation can be selected.
[0155] Robustness: By considering multiple optimization objectives, the risks that may be caused by a single objective are reduced, and the stability and reliability of the system are improved.
[0156] Adaptive Adjustment: As new data continuously flows in, the model can continuously learn and optimize its decision-making process, gradually improving the effect of load balancing.
[0157] Transparency and Controllability: Provide intuitive decision-making support tools for managers, enabling them to flexibly adjust resource allocation according to business priorities.
[0158] Through a carefully designed objective function and an advanced multi-objective optimization algorithm, the present invention not only achieves high-precision server load prediction but also provides a scientific basis for dynamically adjusting resource allocation, which helps to improve the overall operation efficiency and service quality of the data center.
[0159] According to the generated resource allocation plan, the method of the present invention adjusts the load distribution of the server cluster. This may involve specific operations such as virtual machine migration, container rescheduling, load balancer configuration adjustment, etc. Preferably, the present invention adopts a progressive adjustment strategy, that is, only small adjustments are made within each time window (such as 5 minutes). This strategy can reduce the overhead caused by frequent large-scale migrations while maintaining the stability of the system.
[0160] It should be noted that the load balancing strategy of the present invention not only considers the current and predicted load conditions, but also takes into account the effects of historical scheduling decisions. By introducing a reinforcement learning mechanism, the system of the present invention can continuously optimize its decision-making strategy and gradually improve the effect of load balancing.
[0161] Finally, the present invention also proposes a server load balancing prediction system based on deep learning. The system includes the following modules:
[0162] Data acquisition module 1, which is used to acquire the real-time load data and historical load data of the server. This module can collect data through server monitoring agents, log collectors, etc., and perform preliminary data cleaning and formatting processing.
[0163] Topological embedding module 2, which is used to perform topological embedding transformation based on the real-time load data and the historical load data. This module implements the topological embedding function described above, mapping high-dimensional load data to a topological space with geometric meaning.
[0164] Group theory transformation module 3, which is used to perform group theory transformation according to the result of the topological embedding transformation. This module realizes the rotation invariance processing of data through the group action of the special orthogonal group SO(n).
[0165] Spectral decomposition module 4, which is used to perform spectral decomposition based on the result of the group theory transformation. This module performs eigenvalue decomposition on the data to reveal the internal structure of the data.
[0166] Nonlinear transformation module 5, which is used to perform nonlinear transformation according to the result of the spectral decomposition. This module enhances the expression ability of the model by introducing a nonlinear activation function.
[0167] Time series prediction module 6, which is used to perform time series prediction using a long short-term memory network based on the result of the nonlinear transformation. This module is the core of the entire system and is responsible for generating the final load prediction result.
[0168] Output module 7, which is used to output the future load prediction result of the server. This module is not only responsible for the output of the result, but also includes the visualization and interpretation functions of the result, which is convenient for administrators to understand and use the prediction result.
[0169] These modules work closely together to form a complete load prediction and balancing system. Through this modular design, the system of the present invention has good scalability and maintainability. For example, a specific module can be replaced or upgraded as needed without affecting the operation of the overall system.
[0170] Generally speaking, the server load balancing prediction method and system based on deep learning proposed by the present invention achieve high-precision load prediction and intelligent load balancing by innovatively combining topology, group theory, spectral decomposition, and deep learning techniques. This can not only improve the resource utilization efficiency of the data center but also significantly improve the service quality, providing strong technical support for server management in the era of cloud computing and big data.
[0171] To verify the superiority of the server load balancing prediction method and system based on deep learning of the present invention, a set of comparative experiments was designed. The experimental environment is a medium-sized data center, including 100 servers. The load data collection period is 5 minutes, and the prediction time span is 24 hours. The following are the detailed descriptions and experimental results of the examples and comparative examples.
[0172] Example 1: The complete method of the present invention
[0173] In this example, the method proposed by the present invention is fully implemented, including all steps such as topological embedding, group theory transformation, spectral decomposition, nonlinear transformation, and LSTM network prediction. The system parameters are set as follows: Riemannian manifold is used for topological embedding, SO(3) group is used for group theory transformation, the number of LSTM network layers is 3, and the number of hidden units is 128.
[0174] Comparative Example 1: Traditional time series prediction method
[0175] This comparative example uses the ARIMA (Autoregressive Integrated Moving Average) model for load prediction. ARIMA is a classic time series prediction method widely used in various prediction tasks.
[0176] Comparative Example 2: Simple deep learning method
[0177] This comparative example uses a single-layer LSTM network to directly predict the original load data without including the preprocessing steps such as topological embedding and group theory transformation proposed by the present invention. The parameters of the LSTM network are the same as those in Example 1.
[0178] The following three indicators are selected to evaluate the prediction performance:
[0179] 1. MAPE (Mean Absolute Percentage Error): Reflects the average deviation degree between the predicted value and the actual value.
[0180] 2. RMSE (Root Mean Square Error): Measures the average deviation between the predicted value and the actual value.
[0181] 3. Prediction time: The computing time required to complete the 24-hour load prediction.
[0182] The experimental results are shown in Table 1 below:
[0183] Table 1 Comparison table of experimental results of Example 1 with Comparative Example 1 and Comparative Example 2:
[0184] Method MAPE (%) RMSE Prediction time (s) Example 1 3.2 0.045 2.5 Comparative Example 1 8.7 0.112 1.8 Comparative Example 2 5.9 0.078 3.2
[0185] It can be seen from the experimental results in Table 1 that the method of the present invention (Example 1) is significantly superior to the other two methods in terms of prediction accuracy. The specific analysis is as follows:
[0186] 1. MAPE: The method of the present invention controls MAPE at 3.2%, which is much lower than 8.7% of the traditional ARIMA method and 5.9% of the simple LSTM method. This means that the prediction results of the present invention are closer to the actual load situation and can provide a more reliable basis for load balancing decisions.
[0187] 2. RMSE: The RMSE of the method of the present invention is 0.045, also leading far ahead of the other two methods. The lower RMSE indicates that the prediction results of the present invention have smaller fluctuations and higher stability.
[0188] 3. Prediction time: Although the method of the present invention has a higher computational complexity than the traditional ARIMA method, the prediction time is only 0.7 seconds more than that of ARIMA, and at the same time 0.7 seconds faster than the simple LSTM method. Considering the significant improvement in prediction accuracy, this slight increase in time is completely acceptable.
[0189] These results fully prove the effectiveness of the innovative points proposed by the present invention:
[0190] 1. Topological embedding and group theory transformation: These two steps effectively capture the geometric structure and symmetry of the load data, providing a more meaningful feature representation for subsequent prediction. This explains why the method of the present invention can significantly improve the prediction accuracy.
[0191] 2. Spectral decomposition and non-linear transformation: These steps further extract the key features of the data and introduce non-linear expression ability. This enables the method of the present invention to handle complex load patterns, such as periodic fluctuations, sudden peaks, etc.
[0192] 3. Multi-layer LSTM structure: Compared with the simple single-layer LSTM (Comparative Example 2), the multi-layer structure of the present invention can learn more complex time-dependent relationships, thus improving the accuracy of long-term prediction.
[0193] It should be noted that while maintaining high prediction accuracy, the computational efficiency of the method of the present invention also remains at a good level. This is attributed to a series of optimization strategies adopted in the algorithm design, such as using SVD for spectral decomposition and using matrix trace operations for dimensionality reduction, etc.
[0194] Generally speaking, the results of this group of experiments strongly prove the superiority of the present invention in the server load prediction task. Compared with traditional methods and simple deep learning methods, the present invention not only significantly improves the prediction accuracy but also maintains an acceptable computational efficiency. This means that in practical applications, using the method of the present invention can more accurately predict the server load, thereby formulating a better load balancing strategy, and ultimately improving the resource utilization efficiency and service quality of the data center.
[0195] It should be noted that the above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A server load balancing prediction method based on deep learning, characterized in that, Including: An acquisition step, including: acquiring real-time load data and historical load data of a server; a processing step, including: performing a topological embedding transformation based on the real-time load data and the historical load data; performing a group theory transformation according to the result of the topological embedding transformation; performing a spectral decomposition based on the result of the group theory transformation; performing a non-linear transformation according to the result of the spectral decomposition; performing a time series prediction using a long short-term memory network based on the result of the non-linear transformation; an output step, including: outputting a future load prediction result of the server; The topological embedding transformation specifically includes: mapping the server load data onto a Riemannian manifold; using a Riemannian exponential map to achieve the topological embedding of the data; The group theory transformation specifically includes: selecting the special orthogonal group SO(n) as the transformation group; achieving the rotational invariance of the data through group action; The spectral decomposition specifically includes: performing an eigenvalue decomposition on the data after the group theory transformation; obtaining an eigenvector matrix and an eigenvalue diagonal matrix; The non-linear transformation specifically includes: performing a non-linear mapping on the result of the spectral decomposition using an activation function; calculating the trace of the transformed matrix as the output; The long short-term memory network performing the time series prediction specifically includes: processing the data sequence after the non-linear transformation using an LSTM network; outputting the final prediction result through a softmax function; It also includes a model training step: training the long short-term memory network using the historical load data; optimizing the network parameters through a backpropagation algorithm; It also includes a load balancing strategy formulation step: generating a server resource allocation plan based on the future load prediction result; adjusting the load distribution of the server cluster according to the resource allocation plan; The topological embedding transformation adopts the following topological embedding function: , where is an n-dimensional Euclidean space, n is the number of server load characteristics, is a Riemannian manifold, and the specific implementation of this function is as follows: , here is the Riemannian exponential map, is a fixed point on the manifold, is an orthonormal basis in the tangent space, is the -th eigenvalue of the load data.
2. The method according to claim 1, characterized in that, It also includes a data preprocessing step: performing a normalization process on the acquired load data; removing outlier and noise data.
3. A deep learning-based server load balancing prediction system for performing the method according to any one of claims 1-2, characterized in that, Including: A data acquisition module for acquiring the real-time load data and the historical load data of the server; A topological embedding module for performing a topological embedding transformation based on the real-time load data and the historical load data; A group theory transformation module for performing a group theory transformation according to the result of the topological embedding transformation; A spectral decomposition module for performing a spectral decomposition based on the result of the group theory transformation; A non-linear transformation module for performing a non-linear transformation according to the result of the spectral decomposition; A time series prediction module for performing a time series prediction using a long short-term memory network based on the result of the non-linear transformation; An output module for outputting a future load prediction result of the server.
Citation Information
Patent Citations
Group convolution link prediction method based on fusion of multi-objective optimization and evolutionary learning
CN118194930A
Edge cloud computing load balancing method
CN119094531A