Load balancing algorithm based on deep learning

Through the load balancing algorithm based on deep learning, using deep neural networks and reinforcement learning mechanisms, the problems of insufficient intelligence and global optimization of traditional load balancing algorithms in large-scale and dynamically changing load distribution are solved, and efficient and intelligent load balancing is achieved.

CN120045309APending Publication Date: 2025-05-27HAIER CONSUMER FINANCE CO LTD

Patent Information

Application Number
CN202411862425.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-17
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

When traditional load balancing algorithms deal with large-scale and dynamically changing load allocation problems, they lack intelligence and global optimization capabilities, resulting in low load allocation efficiency and extended response time.

Method used

The load balancing algorithm based on deep learning is adopted to learn the time series data of server load through deep neural networks to predict future load changes, and adjust the load allocation strategy online based on the reinforcement learning mechanism to achieve global optimal load balancing.

Benefits of technology

It improves the intelligence and adaptability of load balancing, reduces response time, improves overall resource utilization and user experience, and achieves global optimal load allocation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045309A_ABST
    Figure CN120045309A_ABST
Patent Text Reader

Abstract

The invention relates to a load balancing algorithm based on deep learning. The load balancing algorithm comprises the following steps: step 1, system initialization and model preparation; 2, receiving and preprocessing a user request; step 3, collecting and analyzing server load data; 4, intelligent load distribution decision making; 5, processing and forwarding the request; step 6, an error processing and retry mechanism; 7, evaluating and optimizing the model and the strategy; 8, the user requests are completed, and the system is in a standby state. According to the invention, through the technical means of deep learning modeling, reinforcement learning mechanism, adaptive queue management and global load optimization, intelligent prediction, dynamic adjustment and global optimization of the server load are realized; and the positive effects of improving the load balancing efficiency, shortening the response time, improving the resource utilization rate, improving the user experience and the like are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of cloud computing, and in particular relates to a load balancing algorithm based on deep learning. Background Art

[0002] In cloud computing and data center environments, server clusters need to handle a large number of concurrent requests. Traditional load balancing algorithms, such as polling and random selection, mainly rely on static decision rules or simple feedback mechanisms, and are not very suitable for handling such large-scale, dynamically changing load distribution problems. In recent years, the rapid development of deep learning technology has provided new ideas for solving such problems. Deep learning can effectively learn and simulate complex, dynamically changing system behaviors, and is very suitable for load balancing.

[0003] The main problems faced by existing load balancing technologies include insufficient adaptability to dynamic changes in server instances, resulting in inefficient load distribution and extended response time; traditional algorithms rely on static decision rules and lack a global optimization perspective, making it impossible to achieve optimal cluster throughput and service levels. These problems highlight the shortcomings of traditional load balancing technologies in terms of intelligence, global optimization, and online learning capabilities, and they urgently need to be improved by introducing advanced technical means. Summary of the invention

[0004] (I) Purpose of the invention

[0005] In order to overcome the above shortcomings, the purpose of the present invention is to provide a load balancing algorithm based on deep learning to solve the above technical problems.

[0006] (II) Technical solution

[0007] To achieve the above objectives, the technical solutions provided by this application are as follows:

[0008] A load balancing algorithm based on deep learning includes the following steps:

[0009] Step 1: System initialization and model preparation: Start the load balancing system, perform necessary initialization operations, check the status of the deep learning model, and use historical data for model training and update if necessary.

[0010] Step 2: User request reception and preprocessing: The front-end load balancer receives user requests, performs preliminary identification and classification on the requests, and records logs.

[0011] Step 3: Server load data collection and analysis: deploy monitoring agents to collect key performance indicators of CPU, memory, and bandwidth usage, use tools to regularly pull server performance data, and send it to the deep learning model in real time;

[0012] Step 4: Intelligent load distribution decision, combining deep learning models and real-time server load data to intelligently determine the request distribution plan, and if necessary, execute global optimization strategies to achieve optimal performance of the entire cluster;

[0013] Step 5: Request processing and forwarding: Send the request to the optimal server and record the processing time, hit rate, and error rate. The server processes the request and feeds back the result. The system updates the load data and monitoring indicators based on the result.

[0014] Step 6: Error handling and retry mechanism. For failed requests, the system triggers the retry mechanism and records the failure log for subsequent analysis. If the request processing fails, error handling is performed and attempts are made to reallocate the request.

[0015] Step 7: Model and strategy evaluation and optimization: Regularly evaluate the accuracy of the model and the effectiveness of the strategy to ensure continuous optimization and improvement of system performance, and adjust model parameters and load balancing strategies based on the evaluation results.

[0016] Step 8: User request completed and system on standby. The processing of the user request is completed and the operation log is recorded. The system is ready to receive the next user request, and early warning and pre-resource adjustment are performed for frequently occurring, high-pressure time periods. The system is on standby to process new requests.

[0017] Preferably, in step 2, an Nginx or Apache Web server is used as a front-end load balancer to receive user requests, the requests are recorded by a logger, and preliminary distribution processing is performed, and the request data is encapsulated and sent to a deep learning processing model.

[0018] Preferably, the deep learning model is used to predict the optimal load distribution of the server. Model training needs to be regularly optimized using historical data, collect past request logs and server performance indicator data, select a suitable neural network architecture and configure hyperparameters, use historical data to train the model, so that it learns request patterns and load distribution trends, and the model is implemented using the PyTorch machine learning framework.

[0019] Preferably, the collecting of past request logs and server performance indicator data specifically includes the following steps:

[0020] A1 Determine the data source and clarify the specific location where request logs and performance indicator data are stored on the server, including the server's log file system, specific tables in the database, or data storage locations generated by tools specifically used for performance monitoring;

[0021] A2 Data collection frequency and scope: Determine the frequency of data collection, which depends on the load change speed of the server and the real-time requirements. If the server load changes quickly, a higher collection frequency is required; if the load changes relatively slowly, the collection frequency can be reduced; determine the scope of collected data, including request attributes and server performance indicators. Attributes include request time, source IP, request method, request path and response status code. Server performance indicators include CPU usage, memory usage, disk I / O speed and network bandwidth usage;

[0022] A3 Data storage and management Select an appropriate data storage method to ensure data security and accessibility, use relational databases, non-relational databases or distributed file systems to store data, establish data storage structures and indexes to quickly query and analyze data, back up data regularly to prevent data loss, and use automated backup tools to back up data to local disks, network storage or cloud storage.

[0023] Preferably, selecting a suitable neural network architecture and configuring hyperparameters specifically includes the following steps:

[0024] B1 analysis data characteristics:

[0025] Analyze the collected request logs and server performance indicator data to understand the characteristics and patterns of the data, including observing the time distribution of requests, the changing trends of server performance indicators, and the correlation between different request types.

[0026] Choose the appropriate neural network architecture based on the data characteristics. If the data has spatial correlation, use CNN; if the data has time series characteristics, that is, the request pattern and load distribution change over time, use LSTM;

[0027] B2 selects the neural network architecture:

[0028] If you choose the LSTM architecture, determine the number of LSTM layers and the number of hidden units. The number of LSTM layers and the number of hidden units determine the model's ability to remember and fit time series data. Select appropriate activation functions for the forget gate, input gate, and output gate. Usually, the forget gate and input gate use the Sigmoid function, and the output gate uses the Tanh function. Determine whether to use bidirectional LSTM. Bidirectional LSTM can consider both the forward and reverse information of time series data at the same time, improving the performance of the model, but it will also increase the computational complexity.

[0029] B3 configures hyperparameters,

[0030] Learning rate: Start with a learning rate of 0.001 and then adjust it based on the training of the model;

[0031] Batch size, between 32 and 256.

[0032] Regularization parameter, select the appropriate regularization parameter through cross-validation method;

[0033] Optimization algorithm, choose a suitable optimization algorithm. Different optimization algorithms differ in convergence speed, stability and adaptability to different types of data. Choose a suitable optimization algorithm according to the specific situation.

[0034] Preferably, using historical data to train the model to learn request patterns and load distribution trends specifically includes the following:

[0035] Data preprocessing: preprocessing historical data, including data cleaning, normalization, and standardization operations;

[0036] Divide the data set and divide the preprocessed historical data into training set, validation set and test set. The training set is used to train the model, the validation set is used to adjust the hyperparameters and monitor the model training process, and the test set is used to evaluate the final performance of the model.

[0037] Model training, using the training set data to train the selected neural network architecture. During the training process, the optimization algorithm is used to continuously update the model parameters so that the model's prediction results are as close as possible to the actual load distribution situation.

[0038] Monitor the training process by observing the loss function values ​​and accuracy indicators on the training set and validation set. If the performance of the model on the training set continues to improve, but the performance on the validation set begins to decline, it means that the model is overfitting and corresponding measures need to be taken, including adding regularization terms and reducing model complexity.

[0039] Preferably, step 4 includes a global optimization decision-making module, which infers the current optimal allocation strategy through machine learning and updates the decision model of the load balancing algorithm in real time to adapt to the dynamic changes of the cluster.

[0040] The step 4 includes an online update mechanism of reinforcement learning, using a DQN or PPO reinforcement learning algorithm to update the strategy, receive feedback on the request scheduling results, and adjust the next request allocation to adapt to the changing user request pattern and server status;

[0041] The step 4 includes a multi-queue strategy optimization module, which uses a deep learning model to predict and optimize the request queue management strategy, formulates a sub-queue for each server instance, predicts the waiting time and processing time of different queues through the deep learning model, and dynamically adjusts the request allocation to each sub-queue.

[0042] Preferably, step 5 includes a load balancing intelligent decision-making and forwarding module, which makes intelligent decisions to optimize resource allocation and request forwarding based on a deep learning model and a reinforcement learning strategy, obtains the optimal server allocation result according to the real-time load data of the server and the learning model, implements an intelligent forwarding strategy, and routes the request to the optimal server for processing;

[0043] The step 5 includes a request processing and status feedback module. The server processes the request and feeds back the result. At the same time, the system updates the load status and records the processing result. The request is forwarded to the optimal server and the processing time, hit rate and error rate are recorded. The system updates the load data and monitoring indicators according to the results returned by the server. For failed requests, the system triggers a retry mechanism and records the failure log for subsequent analysis.

[0044] Preferably, in step 7, offline log data is used to test the accuracy of the model, and operation data is collected to analyze the execution effect of the current load balancing strategy.

[0045] Preferably, step 8 also includes recording operation logs and performing subsequent performance analysis and monitoring.

[0046] Beneficial effects:

[0047] 1. Improve the intelligence and adaptability of load balancing: Traditional load balancing algorithms cannot adapt to the dynamic changes of server instances, resulting in low load distribution efficiency and long response time.

[0048] 2. Achieve global optimal load balancing: Requests must be allocated according to the global optimal strategy to improve the throughput and service level of the entire cluster.

[0049] 3. Strengthen the online learning and prediction capabilities of load balancing: dynamically collect and analyze the operating data of service instances, and dynamically adjust the load distribution strategy. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1 It is the overall flow chart of the present invention. DETAILED DESCRIPTION

[0051] In order to make the purpose, technical solutions and advantages of the present invention more clear, the following is a detailed description of the present invention in conjunction with the specific implementation methods and with reference to the attached drawings. Figure 1 , the present invention is further described in detail. It should be understood that these descriptions are exemplary only and are not intended to limit the scope of the present invention. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessary confusion of the concept of the present invention.

[0052] The present invention provides a load balancing algorithm based on deep learning, comprising the following steps:

[0053] Step 1: System initialization and model preparation: Start the load balancing system, perform necessary initialization operations, check the status of the deep learning model, and use historical data for model training and update if necessary.

[0054] Step 2: User request reception and preprocessing: The front-end load balancer receives user requests, performs preliminary identification and classification on the requests, and records logs.

[0055] Step 3: Server load data collection and analysis: deploy monitoring agents to collect key performance indicators of CPU, memory, and bandwidth usage, use tools to regularly pull server performance data, and send it to the deep learning model in real time;

[0056] Step 4: Intelligent load distribution decision, combining deep learning models and real-time server load data to intelligently determine the request distribution plan, and if necessary, execute global optimization strategies to achieve optimal performance of the entire cluster;

[0057] Step 5: Request processing and forwarding: Send the request to the optimal server and record the processing time, hit rate, and error rate. The server processes the request and feeds back the result. The system updates the load data and monitoring indicators based on the result.

[0058] Step 6: Error handling and retry mechanism. For failed requests, the system triggers the retry mechanism and records the failure log for subsequent analysis. If the request processing fails, error handling is performed and attempts are made to reallocate the request.

[0059] Step 7: Model and strategy evaluation and optimization: Regularly evaluate the accuracy of the model and the effectiveness of the strategy to ensure continuous optimization and improvement of system performance, and adjust model parameters and load balancing strategies based on the evaluation results.

[0060] Step 8: User request completed and system on standby. The processing of the user request is completed and the operation log is recorded. The system is ready to receive the next user request, and early warning and pre-resource adjustment are performed for frequently occurring, high-pressure time periods. The system is on standby to process new requests.

[0061] Preferably, in step 2, an Nginx or Apache Web server is used as a front-end load balancer to receive user requests, the requests are recorded by a logger, and preliminary distribution processing is performed, and the request data is encapsulated and sent to a deep learning processing model.

[0062] Preferably, the deep learning model is used to predict the optimal load distribution of the server. Model training needs to be regularly optimized using historical data, collect past request logs and server performance indicator data, select a suitable neural network architecture and configure hyperparameters, use historical data to train the model, so that it learns request patterns and load distribution trends, and the model is implemented using the PyTorch machine learning framework.

[0063] Preferably, the collecting of past request logs and server performance indicator data specifically includes the following steps:

[0064] A1 Determine the data source and clarify the specific location where request logs and performance indicator data are stored on the server, including the server's log file system, specific tables in the database, or data storage locations generated by tools specifically used for performance monitoring;

[0065] If the server uses a web server such as Apache or Nginx, request logs are usually stored in specific log files such as access.log and error.log. For server performance indicators, data can be obtained from the operating system's performance monitoring tools or specialized monitoring software. Performance monitoring tools include top and vmstat on Linux, and specialized monitoring software includes Zabbix and Nagios.

[0066] A2 Data collection frequency and scope: Determine the frequency of data collection, which depends on the load change speed of the server and the real-time requirements. If the server load changes quickly, a higher collection frequency is required; if the load changes relatively slowly, the collection frequency can be reduced; determine the scope of collected data, including request attributes and server performance indicators. Attributes include request time, source IP, request method, request path and response status code. Server performance indicators include CPU usage, memory usage, disk I / O speed and network bandwidth usage;

[0067] For an e-commerce website server, users' shopping request logs can be collected, including information such as request time, user ID, product ID, requested page, and performance indicators such as server CPU usage, memory usage, and network traffic.

[0068] A3 Data storage and management Choose the appropriate data storage method to ensure data security and accessibility. Use relational databases, non-relational databases, or distributed file systems to store data. Establish data storage structures and indexes to quickly query and analyze data. You can create indexes based on fields such as request time and server ID to improve the efficiency of data retrieval. Back up data regularly to prevent data loss. Use automated backup tools to back up data to local disks, network storage, or cloud storage. Store request logs and server performance indicator data in a MongoDB database, and use timestamp fields as indexes to facilitate querying data by time range. At the same time, automatically back up data to an external hard drive or cloud storage every day.

[0069] Preferably, selecting a suitable neural network architecture and configuring hyperparameters specifically includes the following steps:

[0070] B1 analysis data characteristics:

[0071] Analyze the collected request logs and server performance indicator data to understand the characteristics and patterns of the data, including observing the time distribution of requests, the changing trends of server performance indicators, and the correlation between different request types.

[0072] Choose the appropriate neural network architecture based on the data characteristics. If the data has spatial correlation, for example, there is a certain spatial relationship between the performance indicators of different server nodes, use CNN; if the data has time series characteristics, that is, the request pattern and load distribution change over time, use LSTM;

[0073] By analyzing the server performance indicator data, it is found that there is a certain spatial correlation between the CPU usage and memory usage of different server nodes. At this time, the CNN architecture can be selected to capture this spatial feature. If the request log data shows that the number and type of requests change significantly over time, the LSTM architecture can better learn this time series pattern.

[0074] B2 selects the neural network architecture:

[0075] If you choose the LSTM architecture, determine the number of LSTM layers and the number of hidden units. The number of LSTM layers and the number of hidden units determine the model's ability to remember and fit time series data. Generally speaking, increasing the number of LSTM layers and the number of hidden units can improve the performance of the model, but it will also increase the computational complexity and the risk of overfitting.

[0076] Select appropriate activation functions for the forget gate, input gate, and output gate. Usually, the forget gate and input gate use the Sigmoid function, and the output gate uses the Tanh function. Determine whether to use bidirectional LSTM. Bidirectional LSTM can consider both the forward and reverse information of time series data at the same time, improving the performance of the model, but it will also increase the computational complexity.

[0077] B3 configures hyperparameters,

[0078] Learning rate: Start with a learning rate of 0.001 and then adjust it based on the training of the model;

[0079] Batch size, between 32 and 256.

[0080] Regularization parameter, select the appropriate regularization parameter through cross-validation method;

[0081] Optimization algorithm, choose a suitable optimization algorithm. Different optimization algorithms differ in convergence speed, stability and adaptability to different types of data. Choose a suitable optimization algorithm according to the specific situation.

[0082] Preferably, using historical data to train the model to learn request patterns and load distribution trends specifically includes the following:

[0083] Data preprocessing, preprocessing of historical data, including data cleaning, normalization, and standardization operations; data cleaning can remove outliers and erroneous data, and normalization and standardization can convert data into a unified scale to facilitate model training and optimization.

[0084] For example, you can use the mean normalization method to normalize the data to a range of 0 to 1, that is, for each data sample, subtract the mean of all samples and then divide it by the standard deviation of all samples. This makes the data have the same mean and variance, improving the training effect of the model.

[0085] Divide the data set, divide the preprocessed historical data into training set, validation set and test set. The training set is used to train the model, the validation set is used to adjust the hyperparameters and monitor the model training process, and the test set is used to evaluate the final performance of the model. Generally speaking, the data can be divided into training set, validation set and test set in a ratio of 70%, 15%, and 15%. It can also be adjusted according to specific circumstances. For example, if the amount of data is large, the proportion of the training set can be appropriately increased, and the proportion of the validation set and the test set can be reduced.

[0086] Model training uses the training set data to train the selected neural network architecture. During the training process, the optimization algorithm is used to continuously update the model parameters so that the model's prediction results are as close as possible to the actual load distribution. The number of training rounds and the stopping condition can be set. The number of training rounds refers to the number of times the model completes the training set. Generally speaking, the more training rounds, the better the model performance, but it also increases the training time and the risk of overfitting. The stopping condition can be to stop training when the performance of the model on the validation set no longer improves, or to stop training when the training time reaches a certain limit.

[0087] For example, the Adam optimization algorithm can be used, with a learning rate of 0.001, a batch size of 64, and a training round number of 100. During training, an evaluation is performed on the validation set every 10 rounds, and training is stopped if the loss function value on the validation set does not decrease in 5 consecutive rounds.

[0088] Monitor the training process. Monitor the training process of the model by observing the loss function value and accuracy indicators on the training set and validation set. If the performance of the model on the training set continues to improve, but the performance on the validation set begins to decline, it means that the model is overfitting and corresponding measures need to be taken, including adding regularization terms and reducing model complexity. You can use visualization tools such as TensorBoard to monitor the training process of the model and intuitively understand the performance changes and parameter distribution of the model.

[0089] For example, during the training process, you can plot the change curve of the loss function value on the training set and the validation set with the training round. If you find that the loss function value on the validation set starts to rise after a period of time, it means that the model is overfitting. At this time, you can try to increase the L2 regularization term, or reduce the number of layers and hidden units of the neural network to reduce the complexity of the model.

[0090] Preferably, step 4 includes a global optimization decision-making module, which infers the current optimal allocation strategy through machine learning and updates the decision model of the load balancing algorithm in real time to adapt to the dynamic changes of the cluster.

[0091] The step 4 includes an online update mechanism of reinforcement learning, using a DQN or PPO reinforcement learning algorithm to update the strategy, receive feedback on the request scheduling results, and adjust the next request allocation to adapt to the changing user request pattern and server status;

[0092] The step 4 includes a multi-queue strategy optimization module, which uses a deep learning model to predict and optimize the request queue management strategy, formulates a sub-queue for each server instance, predicts the waiting time and processing time of different queues through the deep learning model, and dynamically adjusts the request allocation to each sub-queue.

[0093] Preferably, step 5 includes a load balancing intelligent decision-making and forwarding module, which makes intelligent decisions to optimize resource allocation and request forwarding based on a deep learning model and a reinforcement learning strategy, obtains the optimal server allocation result according to the real-time load data of the server and the learning model, implements an intelligent forwarding strategy, and routes the request to the optimal server for processing;

[0094] The step 5 includes a request processing and status feedback module. The server processes the request and feeds back the result. At the same time, the system updates the load status and records the processing result. The request is forwarded to the optimal server and the processing time, hit rate and error rate are recorded. The system updates the load data and monitoring indicators according to the results returned by the server. For failed requests, the system triggers a retry mechanism and records the failure log for subsequent analysis.

[0095] Preferably, in step 7, offline log data is used to test the accuracy of the model, and operation data is collected to analyze the execution effect of the current load balancing strategy.

[0096] Preferably, step 8 also includes recording operation logs and performing subsequent performance analysis and monitoring.

[0097] Through deep learning modeling, the present invention uses deep neural networks to learn the time series data of server loads, which involves analyzing a large amount of historical data to identify load patterns and trends. This learning mechanism enables the system to predict future load changes and output the optimal load distribution decision accordingly. This not only improves the intelligence and adaptability of load balancing, but also reduces response time and improves overall resource utilization.

[0098] The introduction of reinforcement learning mechanism enables the system to learn the optimal behavior strategy of service instances and clusters online. In this way, the system can continuously adjust its behavior based on real-time feedback to adapt to changing network conditions and server status. This mechanism enhances the flexibility and responsiveness of the system, allowing the load balancing strategy to more accurately meet actual needs, thereby improving the throughput and service level of the entire cluster.

[0099] Adaptive queue management is achieved by building a request queuing model that dynamically optimizes request queue decisions. This means that the system can intelligently adjust queue management policies based on current server load and request characteristics to ensure that requests are processed efficiently and fairly. This dynamic optimization reduces queue delays, improves the efficiency of request processing, and provides users with a smoother experience.

[0100] Global load optimization is achieved by realizing global load information perception and optimization, which enables the system to optimize resource allocation from the perspective of the entire cluster. In this way, the system can not only improve the performance of a single server, but also improve the service performance of the entire cluster, ensuring stable and efficient services under high load conditions. This global optimization strategy helps to maximize resource utilization, reduce operating costs, and improve user experience.

[0101] In summary, the present invention realizes intelligent prediction, dynamic adjustment and global optimization of server load through technical means of deep learning modeling, reinforcement learning mechanism, adaptive queue management and global load optimization, which brings positive effects such as improved load balancing efficiency, shortened response time, improved resource utilization and improved user experience.

[0102] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.

[0103] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A load balancing algorithm based on deep learning, characterized in that: The following steps are involved: Step 1: System initialization and model preparation: Start the load balancing system, perform necessary initialization operations, check the status of the deep learning model, and use historical data for model training and update if necessary. Step 2: User request reception and preprocessing: The front-end load balancer receives user requests, performs preliminary identification and classification on the requests, and records logs. Step 3: Server load data collection and analysis: deploy monitoring agents to collect key performance indicators of CPU, memory, and bandwidth usage, use tools to regularly pull server performance data, and send it to the deep learning model in real time; Step 4: Intelligent load distribution decision, combining deep learning models and real-time server load data to intelligently determine the request distribution plan, and if necessary, execute global optimization strategies to achieve optimal performance of the entire cluster; Step 5: Request processing and forwarding: Send the request to the optimal server and record the processing time, hit rate, and error rate. The server processes the request and feeds back the result. The system updates the load data and monitoring indicators based on the result. Step 6: Error handling and retry mechanism. For failed requests, the system triggers the retry mechanism and records the failure log for subsequent analysis. If the request processing fails, error handling is performed and attempts are made to reallocate the request. Step 7: Model and strategy evaluation and optimization: Regularly evaluate the accuracy of the model and the effectiveness of the strategy to ensure continuous optimization and improvement of system performance, and adjust model parameters and load balancing strategies based on the evaluation results. Step 8: User request completed and system on standby. The processing of the user request is completed and the operation log is recorded. The system is ready to receive the next user request, and early warning and pre-resource adjustment are performed for frequently occurring, high-pressure time periods. The system is on standby to process new requests.

2. A load balancing algorithm based on deep learning according to claim 1, characterized in that: In step 2, the Nginx or Apache Web server is used as a front-end load balancer to receive user requests. The requests are recorded by the logger and initially distributed and processed. The request data is encapsulated and sent to the deep learning processing model.

3. The deep learning-based load balancing algorithm according to claim 1, characterized in that: The deep learning model is used to predict the optimal load distribution of the server. Model training needs to be regularly optimized using historical data, collect past request logs and server performance indicator data, select a suitable neural network architecture and configure hyperparameters, and use historical data to train the model to learn request patterns and load distribution trends. The model is implemented using the PyTorch machine learning framework.

4. The load balancing algorithm based on deep learning according to claim 3, characterized in that: The collection of past request logs and server performance indicator data specifically includes the following steps: A1 Determine the data source and clarify the specific location where request logs and performance indicator data are stored on the server, including the server's log file system, specific tables in the database, or data storage locations generated by tools specifically used for performance monitoring; A2 Data collection frequency and scope: Determine the frequency of data collection, which depends on the load change speed of the server and the real-time requirements. If the server load changes quickly, a higher collection frequency is required; if the load changes relatively slowly, the collection frequency can be reduced; determine the scope of collected data, including request attributes and server performance indicators. Attributes include request time, source IP, request method, request path and response status code. Server performance indicators include CPU usage, memory usage, disk I / O speed and network bandwidth usage; A3 Data storage and management Select an appropriate data storage method to ensure data security and accessibility, use relational databases, non-relational databases or distributed file systems to store data, establish data storage structures and indexes to quickly query and analyze data, back up data regularly to prevent data loss, and use automated backup tools to back up data to local disks, network storage or cloud storage.

5. The deep learning-based load balancing algorithm according to claim 3, characterized in that: Selecting a suitable neural network architecture and configuring hyperparameters specifically includes the following steps: B1 analysis data characteristics: Analyze the collected request logs and server performance indicator data to understand the characteristics and patterns of the data, including observing the time distribution of requests, the changing trends of server performance indicators, and the correlation between different request types. Choose the appropriate neural network architecture based on the data characteristics. If the data has spatial correlation, use CNN; if the data has time series characteristics, that is, the request pattern and load distribution change over time, use LSTM; B2 selects the neural network architecture: If you choose the LSTM architecture, determine the number of LSTM layers and the number of hidden units. The number of LSTM layers and the number of hidden units determine the model's ability to remember and fit time series data. Select appropriate activation functions for the forget gate, input gate, and output gate. Usually, the forget gate and input gate use the Sigmoid function, and the output gate uses the Tanh function. Determine whether to use bidirectional LSTM. Bidirectional LSTM can consider both the forward and reverse information of time series data at the same time, improving the performance of the model, but it will also increase the computational complexity. B3 configures hyperparameters, Learning rate: Start with a learning rate of 0.001 and then adjust it based on the training of the model; Batch size, between 32 and 256. Regularization parameter, select the appropriate regularization parameter through cross-validation method; Optimization algorithm, choose a suitable optimization algorithm. Different optimization algorithms differ in convergence speed, stability and adaptability to different types of data. Choose a suitable optimization algorithm according to the specific situation.

6. The deep learning-based load balancing algorithm according to claim 3, characterized in that: Using historical data to train the model to learn request patterns and load distribution trends includes the following: Data preprocessing: preprocessing historical data, including data cleaning, normalization, and standardization operations; Divide the data set and divide the preprocessed historical data into training set, validation set and test set. The training set is used to train the model, the validation set is used to adjust the hyperparameters and monitor the model training process, and the test set is used to evaluate the final performance of the model. Model training, using the training set data to train the selected neural network architecture. During the training process, the optimization algorithm is used to continuously update the model parameters so that the model's prediction results are as close as possible to the actual load distribution situation. Monitor the training process by observing the loss function values ​​and accuracy indicators on the training set and validation set. If the performance of the model on the training set continues to improve, but the performance on the validation set begins to decline, it means that the model is overfitting and corresponding measures need to be taken, including adding regularization terms and reducing model complexity.

7. The deep learning-based load balancing algorithm according to claim 1, characterized in that: Step 4 includes a global optimization decision-making module, which infers the current optimal allocation strategy through machine learning and updates the decision model of the load balancing algorithm in real time to adapt to the dynamic changes of the cluster. The step 4 includes an online update mechanism of reinforcement learning, using a DQN or PPO reinforcement learning algorithm to update the strategy, receive feedback on the request scheduling results, and adjust the next request allocation to adapt to the changing user request pattern and server status; The step 4 includes a multi-queue strategy optimization module, which uses a deep learning model to predict and optimize the request queue management strategy, formulates a sub-queue for each server instance, predicts the waiting time and processing time of different queues through the deep learning model, and dynamically adjusts the request allocation to each sub-queue.

8. The deep learning-based load balancing algorithm according to claim 1, characterized in that: The step 5 includes a load balancing intelligent decision-making and forwarding module, which makes intelligent decisions to optimize resource allocation and request forwarding based on a deep learning model and a reinforcement learning strategy, obtains the optimal server allocation result according to the real-time server load data and the learning model, implements an intelligent forwarding strategy, and routes the request to the optimal server for processing; The step 5 includes a request processing and status feedback module. The server processes the request and feeds back the result. At the same time, the system updates the load status and records the processing result. The request is forwarded to the optimal server and the processing time, hit rate and error rate are recorded. The system updates the load data and monitoring indicators according to the results returned by the server. For failed requests, the system triggers a retry mechanism and records the failure log for subsequent analysis.

9. The deep learning-based load balancing algorithm according to claim 1, characterized in that: In step 7, offline log data is used to test the accuracy of the model, and operation data is collected to analyze the execution effect of the current load balancing strategy.

10. The deep learning-based load balancing algorithm according to claim 1, characterized in that: The step 8 also includes recording operation logs and performing subsequent performance analysis and monitoring.

Citation Information

Patent Citations

  • Thermal station load prediction and optimization control method and system based on distributed machine learning

    CN118153779A

  • Cloud resource automatic allocation system

    CN118363765A

  • Edge computing load balancing system based on artificial intelligence optimization

    CN118467168A

  • Self-adaptive server load balancing method based on machine learning

    CN118860667A

Cited By

  • High-concurrency lightweight data channel adaptive load balancing method based on large model

    CN120980082A

  • Big model based high concurrency lightweight data channel adaptive load balancing method

    CN120980082B