Rpc load balancing method, server, storage medium and program product
Patent Information
- Application Number
- CN202510285972.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2026-09-11
AI Technical Summary
[0004]本申请提供一种RPC负载均衡方法、服务器、存储介质及程序产品,用以解决分布式系统中RPC请求处理的性能低,影响整个分布式系统的稳定性和性能,进而影响分布式系统的服务质量的问题
Smart Images

Figure CN122733486A_ABST
Abstract
Description
Technical Field
[0001] This application relates to computer technology, and more particularly to an RPC load balancing method, server, storage medium, and program product. Background Technology
[0002] In modern distributed systems, Remote Procedure Call (RPC) is the primary means of communication between clients and servers, and its performance directly affects the overall response speed and stability of the system.
[0003] On the server side, existing RPC load balancing methods mostly employ simple round-robin algorithms or weighted round-robin algorithms based on static weights to allocate worker threads for RPC requests. Under high concurrency or load fluctuations, some worker threads may experience performance degradation due to overload, affecting the stability and performance of the entire distributed system, and consequently impacting the service quality of the distributed system. Summary of the Invention
[0004] This application provides an RPC load balancing method, server, storage medium, and program product to solve the problem of low performance in RPC request processing in distributed systems, which affects the stability and performance of the entire distributed system and thus the service quality of the distributed system.
[0005] Firstly, this application provides an RPC load balancing method, including:
[0006] In response to receiving an RPC request, the current allocation weight of each worker thread is obtained, wherein the allocation weight of each worker thread is dynamically adjusted based on the real-time load information of each worker thread and the predicted load information at future times;
[0007] Worker threads are assigned to the RPC request based on the current allocation weight of each worker thread.
[0008] Secondly, this application provides an RPC load balancing method applied to a distributed storage server, the method comprising:
[0009] Receive RPC requests sent by client devices, wherein the RPC requests include at least one of the following: read data request, write data request, disk management request, snapshot management request, data migration request, and cluster management request;
[0010] Obtain the current allocation weight of each worker thread, wherein the allocation weight of each worker thread is dynamically adjusted based on the real-time load information of each worker thread and the predicted load information at future times;
[0011] Worker threads are assigned to the RPC request based on the current allocation weight of each worker thread.
[0012] Thirdly, this application provides a server, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to cause the server to perform the methods provided in any of the foregoing aspects.
[0013] Fourthly, this application provides a computer-readable storage medium storing computer-executable instructions that, when executed by a processor, implement the method provided in any of the foregoing aspects.
[0014] Fifthly, this application provides a computer program product, including a computer program that, when executed by a processor, implements the methods provided in any of the foregoing aspects.
[0015] The RPC load balancing method, server, storage medium, and program products provided in this application dynamically adjust the allocation weight of each worker thread based on its real-time load information and predicted load information at future times. This approach considers not only the real-time load information but also the predicted load information at future times, improving the rationality and accuracy of the allocation weights. Furthermore, the control device allocates worker threads to RPC requests based on their current allocation weights. Under high concurrency or load fluctuations, this dynamically adapts to the real-time load of worker threads, preventing some threads from being overloaded while others are idle. This ensures that each worker thread operates under optimal load, improving the stability and performance of the distributed system, and ultimately enhancing the service quality of the distributed system. Attached Figure Description
[0016] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0017] Figure 1 A flowchart of an example embodiment of the RPC load balancing method provided in this application;
[0018] Figure 2 A flowchart illustrating a method for dynamically adjusting the weights assigned to worker threads, as provided in an exemplary embodiment of this application;
[0019] Figure 3 A flowchart of RPC load balancing provided as another exemplary embodiment of this application;
[0020] Figure 4A framework diagram of a control system provided for an exemplary embodiment of this application;
[0021] Figure 5 A flowchart illustrating the processing of an RPC request provided for an exemplary embodiment of this application;
[0022] Figure 6 A flowchart illustrating the dynamic adjustment of worker thread weights as provided in an exemplary embodiment of this application;
[0023] Figure 7 A flowchart of an RPC load balancing method for a distributed storage system provided in an exemplary embodiment of this application;
[0024] Figure 8 This is a structural block diagram of a computing device according to an embodiment of this application;
[0025] Figure 9 This is a schematic diagram of the structure of a server provided in an embodiment of this application.
[0026] The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to particular embodiments. Detailed Implementation
[0027] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0028] It should be noted that the user information (including but not limited to user device information, user attribute information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with relevant laws, regulations and standards, and corresponding operation entry points are provided for users to choose to authorize or refuse.
[0029] First, let me explain the terms used in this application:
[0030] Remote Procedure Call (RPC): A protocol that allows programs to communicate between processes on different computers by calling functions or methods of remote services to implement function calls in a distributed system.
[0031] The management and control system (also known as RiverMaster) is the core management and control system of the Elastic Block Storage (EBS) system. It is mainly responsible for the management of cloud disk block devices within a certain area (such as the availability zone).
[0032] Load balancing (LB) is a technique that effectively distributes workloads (such as requests and tasks) across multiple computing resources (such as servers and threads) to optimize resource utilization, maximize throughput, minimize response time, and ensure system stability and high availability.
[0033] In distributed systems, traditional RPC load balancing schemes mainly include methods such as Round Robin, Least Connections, Weighted Round Robin, Least Load, and Consistent Hashing. These methods achieve request distribution and load balancing to a certain extent, but they still have many shortcomings in dealing with high concurrency and dynamic load environments.
[0034] The polling algorithm distributes requests to worker threads sequentially in a loop, making it simple and easy to implement. However, this method ignores the actual load on each worker thread and cannot dynamically adapt to the real-time performance of the threads. When some threads have weak processing capabilities or high loads, they will still continue to receive requests, causing these threads to be overloaded while other threads' resources are not fully utilized, thus affecting the overall response speed and stability of the system.
[0035] The least connections algorithm distributes requests to the thread with the fewest current connections, effectively balancing the load. However, this method makes decisions based solely on the current number of connections, failing to consider thread processing capabilities and task complexity. In high-concurrency environments, frequent changes in the number of connections can lead to unstable load balancing decisions, increased system monitoring overhead, and potential lock contention issues, impacting system performance.
[0036] Weighted round-robin assigns different weights to each thread based on round-robin to reflect differences in their processing capabilities. While more flexible than regular round-robin, the weights are typically statically configured. When system load fluctuates significantly, weighted round-robin cannot respond promptly, leading to uneven resource utilization and degraded system performance.
[0037] Consistent hashing is primarily used in distributed caching and database systems. By mapping requests and threads to a hash ring, it ensures that only a small number of requests need to be reallocated when nodes change. While this method excels in scalability and fault tolerance, its adaptability to load balancing is limited, especially in situations of uneven load distribution or significant differences in thread performance, making it difficult to achieve balanced load allocation.
[0038] In summary, the aforementioned RPC load balancing solutions have their own advantages and disadvantages in different scenarios, but they generally suffer from insufficient adaptability in high-concurrency and dynamic load environments. Under high concurrency or load fluctuations, some worker threads may experience performance degradation due to overload, affecting the stability and performance of the entire distributed system, and consequently impacting the service quality of the distributed system.
[0039] This application provides an RPC load balancing method that dynamically adjusts the weight of each worker thread based on its real-time load information and predicted load information at future times. Under high concurrency or load fluctuations, it dynamically adapts to the real-time load of worker threads, preventing some threads from being overloaded while others are idle, ensuring that each worker thread runs under optimal load, improving the stability and performance of the distributed system, and thus improving the service quality of the system.
[0040] The solution in this embodiment can be applied to RPC load balancing in various distributed systems, such as distributed storage systems and elastic block storage (EBS) systems, without any specific limitations.
[0041] The technical solution of this application and how the technical solution of this application solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will now be described with reference to the accompanying drawings.
[0042] Figure 1 This is a flowchart illustrating an example embodiment of the RPC load balancing method provided in this application. The execution entity in this embodiment is a management and control device used in a distributed system to implement RPC load balancing, such as a management and control server in a distributed storage system. Figure 1 As shown, the specific steps of this method are as follows:
[0043] Step S101: In response to receiving an RPC request, obtain the current allocation weight of each worker thread, wherein the allocation weight of each worker thread is dynamically adjusted based on the real-time load information of each worker thread and the predicted load information at future moments.
[0044] RPC requests can be RPC requests sent by clients or upper-layer applications in a distributed system to the management and control devices of the distributed system, and can be used to implement various functions in the distributed system.
[0045] For example, in a distributed storage system, an RPC request may specifically include at least one of the following: read data request, write data request, disk management request, snapshot management request, data migration request, and cluster management request. Disk management requests include various requests supported by disk management, such as creating and deleting disks. Snapshot management requests include various snapshot-related requests, such as creating and deleting snapshots. Cluster management requests include various cluster management-related requests, such as scaling up and scaling down.
[0046] For a received RPC request, the control device needs to allocate a worker thread to process the RPC request.
[0047] In this embodiment, the control device dynamically (i.e., continuously at multiple different times) adjusts the allocation weight of each worker thread based on its real-time load information and predicted load information at future times. This not only considers the real-time load information of the worker threads but also incorporates their predicted load information at future times, improving the rationality and accuracy of the allocation weights. Specifically, the lower the real-time load and predicted load of a worker thread, the greater its allocation weight.
[0048] The future moment is a concept corresponding to the current moment. It is a moment with a preset duration after the current moment. When the current moment changes, the corresponding future moment will also change accordingly.
[0049] For example, let t represent any current time, and the future time can be represented as t+Δt, where Δt is a preset duration that can be set and adjusted according to actual application needs. For example, Δt can take values of 15 seconds, 10 seconds, etc., without specific limitations here.
[0050] When assigning worker threads to RPC requests, the control device first obtains the current assignment weight of each worker thread.
[0051] Step S102: Assign worker threads to RPC requests according to the current allocation weight of each worker thread.
[0052] In this step, the control device can prioritize assigning RPC requests to worker threads with higher allocation weights based on the current allocation weights of each worker thread.
[0053] For example, the control device can use a weighted round-robin algorithm to prioritize assigning RPC requests to worker threads with higher allocation weights based on the current allocation weights of each worker thread.
[0054] In this embodiment, within a distributed system, the management and control device dynamically adjusts the allocation weights of each worker thread based on its real-time load information and predicted load information for future moments. This approach considers both real-time and predicted load information, improving the rationality and accuracy of the allocation weights. Furthermore, the management and control device assigns worker threads to RPC requests based on their current allocation weights. Under high concurrency or load fluctuations, this dynamically adapts to the real-time load of worker threads, preventing some threads from being overloaded while others remain idle. This ensures that each worker thread operates under optimal load, improving the stability and performance of the distributed system, and ultimately enhancing its service quality.
[0055] Figure 2 This is a flowchart illustrating a method for dynamically adjusting the weight allocation of worker threads, provided as an exemplary embodiment of this application. Building upon the foregoing embodiments, in this embodiment, the control device can periodically adjust the weight allocation of each worker thread based on its real-time load information and predicted load information for future times. For example... Figure 2 As shown, the steps for dynamically adjusting the weights assigned to worker threads are as follows:
[0056] Step S201: At each first time interval, predict the predicted load information of each worker thread at future times based on the real-time load information of each worker thread.
[0057] The first duration is the interval for adjusting the weights assigned to worker threads. A shorter first duration results in a higher frequency of dynamic adjustments to worker thread weights, while a longer first duration results in a lower frequency of dynamic adjustments. The first duration can be set and adjusted according to the actual application scenario requirements; for example, it can be set to 5 seconds, 1 second, etc., without specific limitations here.
[0058] The future moment is a concept corresponding to the current moment. It is a moment after the current moment with a preset duration. When the current moment changes, its corresponding future moment will also change accordingly. For example, let t represent any current moment, and the future moment can be represented as t + Δt, where Δt is the preset duration. It can be set and adjusted according to actual application needs. For example, Δt can take values of 15 seconds, 10 seconds, etc., without specific limitations here.
[0059] In this step, at each first time interval, the control device predicts the future load information of each worker thread based on the real-time load information of each worker thread, and adjusts the allocation weight of each worker thread based on the real-time load information and the predicted load information at the future time. The real-time load information of the worker threads can be the most recently collected real-time load information of the worker threads.
[0060] For example, taking a first duration of 5 seconds as an example, the weight of each worker thread is adjusted every 5 seconds. This allows for timely adjustment of the weight of each worker thread based on its real-time load at the current moment and its predicted load at future moments, thereby improving the accuracy and rationality of the weight allocation.
[0061] In this embodiment, the control device can collect the load information (referred to as real-time load information) of each worker thread in real time (e.g., every 5 seconds or every 15 seconds) and store the real-time load information and collection time of each worker thread. The real-time load information of the worker thread includes at least one of the following: CPU utilization, average response time, and task queue length.
[0062] In one optional embodiment, a pre-trained load prediction model can be used to predict the predicted load information of each worker thread at future times based on the real-time load information of each worker thread. Specifically, the real-time load information of each worker thread, the current time t corresponding to the real-time load information, and the future time (t+Δt) are input into the load prediction model for prediction to obtain the predicted load information of each worker thread at future times.
[0063] The load prediction model is trained from a machine learning model. During use, the control equipment can continuously train and optimize the load prediction model based on newly generated load information.
[0064] Specifically, at every second time interval, historical load information for each worker thread is acquired; based on this historical load information, a load prediction model is trained. The historical load information includes the actual load information of each worker thread at multiple historical moments, with an interval of Δt between adjacent historical moments. If the historical load information is collected at a high frequency, meaning the interval between two adjacent collection moments in the historical load information is less than Δt, then a portion of the historical load information can be sampled, ensuring that the interval between two adjacent historical moments in the sampled historical load information is Δt.
[0065] The second duration is the interval for training and optimizing the load prediction model. A shorter second duration results in a higher frequency of training and optimization of the load prediction model, while a longer second duration results in a lower frequency of training and optimization. The second duration can be set and adjusted according to the actual application scenario requirements. For example, the second duration can be set to 30 minutes, 10 minutes, or several hours, etc., without specific limitations here.
[0066] Specifically, the iterative training / optimization process of the load prediction model is as follows: The actual load information of each worker thread at the previous time step, along with the data from the previous and next time steps, are input into the load prediction model for prediction, resulting in the predicted load information for each worker thread at the next time step. Further, the parameters of the load prediction model are adjusted based on the actual load information and the predicted load information for each worker thread at the next time step.
[0067] For example, based on the actual load information and the predicted load information of each worker thread at the next time step, the cross-entropy loss function value is calculated, and backpropagation is performed based on the cross-entropy loss function value to adjust the parameters of the load prediction model. The training objective is to minimize the difference between the actual load information and the predicted load information of each worker thread at the next time step.
[0068] After multiple iterations of training, the load prediction model that has been trained / optimized is obtained.
[0069] It should be noted that the training strategy used to train the load prediction model, including the learning rate and optimization algorithm, can be designed and adjusted according to the actual application requirements, and no specific limitations are made here.
[0070] In this embodiment, by utilizing a load prediction model to predict the future load of worker threads based on their real-time load information, accurate prediction of future worker thread load can be achieved. Furthermore, by training and optimizing the load prediction model every second time interval based on the historical load information of each worker thread, the prediction accuracy of the load prediction model can be further improved.
[0071] Step S202: Adjust the allocation weight of each worker thread based on the real-time load information of each worker thread and the predicted load information of each worker thread at future time.
[0072] After obtaining the real-time load information of each worker thread and the predicted load information of each worker thread at future time, the control device calculates the new weight of each worker thread based on the real-time load information and the predicted load information of each worker thread at future time, and adjusts the assigned weight of each worker thread to the newly calculated weight.
[0073] Specifically, for any worker thread, the control device calculates the first weight of the worker thread based on the real-time load information and the predicted load information at future times; normalizes the first weight of each worker thread to obtain the second weight of each worker thread; and adjusts the allocation weight of each worker thread according to the second weight of each worker thread.
[0074] Optionally, when calculating the first weight of a worker thread, the control device calculates the predicted load value of the worker thread based on the predicted load information of the worker thread at future times; calculates the real-time load value of the worker thread based on the real-time load information of the worker thread; and calculates the first weight of the worker thread based on the predicted load value and the real-time load value. The first weight of the worker thread is determined by considering the real-time load and predicted load of each worker thread individually, and is also called the single-thread weight of the worker thread.
[0075] The predicted load value for worker threads can be calculated as follows: Predicted load value = a1 × (1 - predicted CPU utilization) + a2 × 1 / predicted average response time + a3 × 1 / (predicted task queue length + 1). Here, a1, a2, and a3 are the normalization coefficients corresponding to CPU utilization, average response time, and task queue length in the predicted load information, respectively, to ensure that the predicted load value for worker threads is normalized to a value within the (0-1) range. The specific values of a1, a2, and a3 can be set according to actual application requirements and are not specifically limited here. For example, a1, a2, and a3 can all be set to 0.3.
[0076] The real-time load value of a worker thread can be calculated as follows: Real-time load value = b1 × (1 - real-time CPU utilization) + b2 × 1 / real-time average response time + b3 × 1 / (real-time task queue length + 1). Here, b1, b2, and b3 are the normalization coefficients corresponding to CPU utilization, average response time, and task queue length in the real-time load information, respectively, to ensure that the real-time load value of the worker thread is normalized to a value within the (0-1) range. The specific values of b1, b2, and b3 can be set according to actual application requirements and are not specifically limited here. For example, b1, b2, and b3 can be set to 0.3, 0.3, and 0.2, respectively.
[0077] The first weight of a worker thread can be calculated as follows: First weight = c1 × real-time load value + c2 × (1 - predicted load value). Here, c1 and c2 are weighting coefficients, which can be set according to actual application requirements; no specific limitation is made here. For example, c1 and c2 can be set to 1 and 0.2 respectively.
[0078] Furthermore, the first weight (i.e., single-thread weight) of each worker thread is normalized to obtain the second weight (also known as normalized weight) of each worker thread.
[0079] For example, the sum of the second weights of each worker thread is calculated to obtain the total weight of each worker thread, and the proportion of the second weight of each worker thread in the total weight is calculated to obtain the second weight of each worker thread.
[0080] Furthermore, the second weight of each worker thread is used as the allocation weight for each worker thread. Alternatively, the decimal part (if any) of the second weight of each worker thread is rounded off and used as the allocation weight for each worker thread.
[0081] It should be noted that the specific calculation formula for allocating weights to each worker thread, based on the real-time load information and the predicted load information of each worker thread at future moments, can be set according to actual application requirements, and is not specifically limited here.
[0082] In this embodiment, at each first time interval, the predicted load information of each worker thread at future times is predicted based on the real-time load information of each worker thread. Based on the real-time load information and the predicted load information of each worker thread at future times, the allocation weight of each worker thread is adjusted. By dynamically adjusting the allocation weight of each worker thread based on its real-time load and predicted load, the rationality and accuracy of the allocation weight can be improved. This allows the RPC load balancing scheme to dynamically adapt to the real-time load of worker threads, preventing some threads from being overloaded while others are idle, ensuring that each worker thread runs under optimal load, improving the stability and performance of the distributed system, and thus improving the service quality of the distributed system.
[0083] Figure 3 A flowchart of RPC load balancing provided for another exemplary embodiment of this application. (e.g.) Figure 3 As shown, the detailed process of RPC load balancing is as follows:
[0084] Step S300: Collect the real-time load information of each working thread in real time, and store the real-time load information and collection time of each working thread.
[0085] This step implements thread monitoring. For details on the implementation principle, please refer to the relevant content in the aforementioned embodiments, which will not be repeated here.
[0086] Step S301: At every second time interval, train the load prediction model based on the historical load information of each worker thread.
[0087] This step enables load prediction for worker threads. For the specific implementation principle, please refer to the relevant content in the aforementioned embodiments, which will not be repeated here.
[0088] Step S302: At each first time interval, predict the predicted load information of each worker thread at future times based on the real-time load information of each worker thread.
[0089] Step S303: Adjust the allocation weight of each worker thread based on the real-time load information of each worker thread and the predicted load information of each worker thread at future time.
[0090] Steps S302-S303 implement dynamic adjustment of the allocation weight of worker threads. For the specific implementation principle, please refer to the relevant content of the aforementioned embodiments, which will not be repeated here.
[0091] Step S304: In response to receiving an RPC request, obtain the current allocation weight of each worker thread.
[0092] Step S305: Assign worker threads to RPC requests based on the current allocation weight of each worker thread.
[0093] Step S306: Add the RPC request to the task queue of the worker thread assigned to the RPC request.
[0094] Step S307: Process the RPC requests in the task queue sequentially using the worker thread allocated to the RPC request, and obtain the processing result of the RPC request.
[0095] Step S308: Output the processing result of the RPC request.
[0096] Steps S304-S308 implement the RPC request processing flow. For the specific implementation principle, please refer to the relevant content of the aforementioned embodiments, which will not be repeated here.
[0097] For the specific implementation principles and technical effects of this embodiment, please refer to the relevant content of the foregoing embodiments, which will not be repeated here.
[0098] Figure 4 This is a framework diagram of a control system provided for an exemplary embodiment of this application. Figure 4 As shown, the management and control system can be divided into an interface layer, a middleware layer, a business layer, a call layer, a data layer, and a monitoring module.
[0099] The interface layer is the RPC server used by the management system to connect with clients. It provides an IO thread pool with multiple IO (input / output) threads, which are responsible for receiving RPC requests from clients.
[0100] The middleware layer includes an RPC scheduling module (i.e., the RPC scheduler), a load balancing module, and a load prediction module. The load prediction module predicts the future load of worker threads. The load balancing module predicts the future load of worker threads based on their real-time load information and dynamically adjusts the allocation weights of worker threads according to both real-time and predicted load information. The RPC scheduling module receives and allocates RPC requests, specifically assigning worker threads to each request based on their allocation weights.
[0101] The business layer provides a worker thread pool containing multiple worker threads, as well as the implementation logic for various functional modules of the distributed system, including disk management, snapshot management, migration management, and cluster management. Worker threads can run the implementation logic to achieve the corresponding functions.
[0102] The call layer provides subsystem call management, which implements the logic of various functional modules in the business layer through subsystem calls.
[0103] The data layer provides data caching and database Data Access Objects (DAO) components. DAO components encapsulate the operational details of database access and provide the business layer with a series of methods for performing CRUD operations on data through interfaces.
[0104] The monitoring module is responsible for implementing functions such as thread monitoring, business monitoring, and database monitoring. Through the thread monitoring function, the load information of each worker thread can be collected in real time, such as the CPU utilization, average response time, and task queue length of the worker thread.
[0105] Figure 5 This is a flowchart illustrating the RPC request processing provided in this embodiment. Based on... Figure 4 The control system shown is as follows: Figure 5 As shown, the processing flow of an RPC request is as follows:
[0106] S50, The client initiates an RPC request to the IO thread;
[0107] S51, the IO thread forwards the RPC request to the RPC scheduler;
[0108] S52, The RPC scheduler requests the load balancing module to make a load balancing decision in order to determine the target thread to be allocated to the RPC request;
[0109] S53. The load balancing module determines the target thread for the RPC request based on the current allocation weight of the worker thread.
[0110] In this embodiment, the load balancing module is also used to dynamically adjust the allocation weight of each worker thread based on the real-time load information of each worker thread and the predicted load information at future times. For specific implementation principles and technical effects, please refer to the relevant content of the foregoing embodiments, which will not be repeated here.
[0111] S54. The load balancing module returns information about the target thread to the RPC scheduler.
[0112] S55, The RPC scheduler assigns RPC requests to target threads in the worker thread pool;
[0113] S56. The target thread processes the RPC request;
[0114] S57. The target thread returns the processing result of the RPC request to the RPC scheduler;
[0115] S58. The RPC scheduler returns the processing result of the RPC request to the IO thread;
[0116] S59. The IO thread returns the processing result of the RPC request to the client.
[0117] based on Figure 4 The control system shown is as follows: Figure 6 As shown, the dynamic adjustment process for worker thread weight allocation is as follows:
[0118] S61. The worker threads in the worker thread pool report real-time load information to the thread monitoring module.
[0119] S62. The thread monitoring module provides the real-time load information reported by the worker threads to the load balancing module.
[0120] The load balancing module can also return a response message (such as a confirmation message) to the thread monitoring module when it receives real-time load information provided by the thread monitoring module.
[0121] S63. The load balancing module predicts the load information of the worker threads at future moments based on the real-time load information of the worker threads.
[0122] S64. The load balancing module dynamically adjusts the allocation weight of the worker threads based on the real-time load information and predicted load information of the worker threads.
[0123] In this embodiment, within a distributed system, the management and control device can dynamically adjust the allocation weights of each worker thread by combining real-time load information and predicted load information for future moments. This approach considers not only real-time load information but also predicted load information for future moments, improving the rationality and accuracy of the allocation weights. Furthermore, the management and control device allocates worker threads to RPC requests based on the current allocation weights of each worker thread. Under high concurrency or load fluctuations, it can dynamically adapt to the real-time load of worker threads, preventing some threads from being overloaded while others are idle. This ensures that each worker thread runs under optimal load, improving the stability and performance of the distributed system, and ultimately enhancing the service quality of the distributed system.
[0124] Figure 7 A flowchart of an RPC load balancing method for a distributed storage system provided as an exemplary embodiment of this application.
[0125] like Figure 7 As shown, the steps for RPC load balancing in a distributed storage system are as follows:
[0126] Step S701: Receive an RPC request sent by the client device, wherein the RPC request includes at least one of the following: read data request, write data request, disk management request, snapshot management request, data migration request, and cluster management request.
[0127] For example, when a client needs to read data stored in a distributed block storage system, it sends an RPC request to the management device of the distributed storage system. These RPC requests typically contain the block address and size of the data to be read.
[0128] For example, when a client needs to write data to a distributed block storage system, it sends an RPC request to the management device of the distributed storage system. These RPC requests typically contain the address and size of the block to be written, and the data to be written itself.
[0129] For example, distributed storage systems typically maintain metadata such as volume information, block mappings, and snapshots. Clients may need to query or modify this metadata. For instance, a client might want to retrieve the current size, remaining space, or block mapping information of a volume, or it might want to create a new snapshot. These operations all require sending RPC requests to the management device of the distributed storage system.
[0130] For example, a client may need to perform management operations on storage volumes, such as creating new volumes, deleting volumes, expanding or shrinking volumes. These operations typically involve the allocation and release of underlying storage resources, thus requiring the client to communicate with the management device of the distributed storage system via RPC requests.
[0131] Step S702: Obtain the current allocation weight of each worker thread, wherein the allocation weight of each worker thread is dynamically adjusted based on the real-time load information of each worker thread and the predicted load information at future moments.
[0132] In this embodiment, the allocation weight of each worker thread is dynamically adjusted periodically based on the real-time load information of each worker thread and the predicted load information at future times.
[0133] Specifically, at each first time interval, based on the real-time load information of each worker thread, the predicted load information of each worker thread at a future time is predicted; based on the real-time load information of each worker thread and the predicted load information of each worker thread at a future time, the allocation weight of each worker thread is adjusted.
[0134] Step S703: Assign worker threads to RPC requests according to the current allocation weight of each worker thread.
[0135] Furthermore, the RPC request is processed by a worker thread allocated to it, the processing result of the RPC request is obtained, and the processing result of the RPC request is returned to the client device.
[0136] For the specific implementation principles and technical effects of this embodiment, please refer to the relevant content of the foregoing embodiments, which will not be repeated here.
[0137] This application provides an adaptive RPC load balancing scheme based on a load prediction model (machine learning model) and real-time thread monitoring. By collecting real-time load information of each worker thread (including but not limited to CPU utilization, average response time, task queue length, and other performance indicators), the machine learning model predicts the future load information of the worker threads. Combining the real-time load and future load of the worker threads, the allocation weight of the threads is dynamically adjusted to achieve intelligent load balancing. This not only improves the system's response speed and stability, but also optimizes resource utilization through predictive adjustments, ensuring that the distributed storage system still performs well under high concurrency and variable loads.
[0138] Figure 8 This is a structural block diagram of a computing device according to an embodiment of this application. Figure 8 As shown, the computing device may include one or more (only one is shown in the figure) processors 801 and memory 802. The memory 802 stores computer programs / instructions, and the processor 801 executes the computer programs / instructions. When the computer programs / instructions are executed by the processor 801, they implement the technical solutions provided in any of the aforementioned method embodiments. Their specific functions and the technical effects they can achieve are similar and will not be repeated here.
[0139] The aforementioned computing device can be understood as an integrated smart terminal, including but not limited to servers, desktop computers, PCs (Personal Computers), all-in-one model machines, mobile phones, tablet computers, or other portable smart terminals, and the computing device may have the model in the above embodiments of this application pre-installed.
[0140] Specifically, this computing device can pre-install various types of models, including but not limited to models in natural language processing, visual processing, speech processing, code processing, and multimodal task processing, thus providing diverse model selection. In different product forms, this computing device can support one or more model usage methods, including but not limited to model training, model invocation, model fine-tuning, model deployment, model inference, and application. In some product forms, this computing device also supports model management, including but not limited to multi-type model management (supporting the management of discriminative, generative, and other types of models), model version control (supporting the control of different model versions), and model evaluation (evaluating model performance and effectiveness based on model evaluation tools). In other product forms, this computing device can also create applications based on models, providing API (Application Programming Interface) calling capabilities. Users can call models into created applications through the API interface, and application management tools are also provided to manage and monitor the applications.
[0141] Furthermore, the computing device may also include data management (supporting the creation and management of model tuning datasets), a training center (providing abundant training resources to help users learn and master AI technology), and basic control capabilities (providing enterprise-level basic control capabilities to ensure the security and efficient operation of the system). Through the above functions, it provides a comprehensive and integrated device for AI development, training, deployment, and application.
[0142] Figure 9 This is a schematic diagram of the structure of a server provided in an embodiment of this application. Figure 9 As shown, the server includes a memory 901 and a processor 902. The memory 901 stores computer-executable instructions and can be configured to store various other data to support operations on the server. The processor 902 is communicatively connected to the memory 901 and executes the computer-executable instructions stored in the memory 901 to implement the technical solutions provided in any of the above method embodiments. Their specific functions and the technical effects they achieve are similar and will not be repeated here.
[0143] Optional, such as Figure 9 As shown, the server also includes other components such as a firewall 903, a load balancer 904, a communication component 905, and a power supply component 906. Figure 9 The diagram only shows a portion of the components and does not imply that the server only includes... Figure 9 The components shown. Figure 9 This example uses a cloud server deployed in the cloud as an example, but the server can also be deployed locally. This embodiment does not make any specific limitations here.
[0144] This application also provides a computer-readable storage medium storing computer-executable instructions. When a processor executes the computer-executable instructions, it implements the method of any of the foregoing embodiments. The specific functions and technical effects to be achieved are not described here.
[0145] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the method of any of the foregoing embodiments. The computer program is stored in a readable storage medium, and at least one processor of the server can read the computer program from the readable storage medium. The execution of the computer program by the at least one processor causes the server to perform the technical solution provided in any of the above method embodiments. The specific functions and the technical effects that can be achieved are not described here.
[0146] This application provides a chip, including a processing module and a communication interface. The processing module is capable of executing the technical solution of the server in the aforementioned method embodiments. Optionally, the chip further includes a storage module (e.g., a memory), which stores instructions. The processing module executes the instructions stored in the storage module, and the execution of the instructions stored in the storage module causes the processing module to execute the technical solution provided in any of the aforementioned method embodiments.
[0147] The integrated modules described above, implemented as software functional modules, can be stored in a computer-readable storage medium. These software functional modules, stored in a storage medium, include several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute some steps of the methods of the various embodiments of this application.
[0148] It should be understood that the aforementioned processor can be a Central Processing Unit (CPU), a graphics processing unit (GPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. A general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in the application can be directly manifested as being executed by a hardware processor, or executed by a combination of hardware and software modules in at least one processor.
[0149] The memory may include high-speed random access memory (RAM) and may also include non-volatile memory, such as at least one disk storage device, and may also be a USB flash drive, external hard drive, read-only memory, disk or optical disc, etc.
[0150] The aforementioned storage device can be object storage service (OSS).
[0151] The aforementioned memory can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Read Only Memory (PROM), Read Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0152] The aforementioned communication components are configured to facilitate wired or wireless communication between the device containing the communication components and other devices. The device containing the communication components can access wireless networks based on communication standards, such as mobile hotspots (WiFi), second-generation (2G), third-generation (3G), fourth-generation (4G) / Long Term Evolution (LTE), fifth-generation (5G), or combinations thereof. In one exemplary embodiment, the communication components receive broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, the communication components also include a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module may be based on Radio Frequency Identification (RFID), infrared, Ultra Wide Band (UWB), Bluetooth, and other technologies.
[0153] The aforementioned power supply components provide power to various components within the device in which they reside. These power supply components may include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power to the device in which they reside.
[0154] The aforementioned storage medium can be implemented from any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The storage medium can be any available medium accessible to general-purpose or special-purpose computers.
[0155] An exemplary storage medium is coupled to a processor, enabling the processor to read information from and write information to the storage medium. Alternatively, the storage medium can be an integral part of the processor. The processor and storage medium can reside within an application-specific integrated circuit (ASIC). Alternatively, the processor and storage medium can exist as discrete components within an electronic device or host device.
[0156] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0157] The order of the embodiments described above is merely for illustrative purposes and does not represent the superiority or inferiority of the embodiments. Furthermore, some processes described in the above embodiments and accompanying drawings include multiple operations appearing in a specific order. However, it should be clearly understood that these operations may not be executed in the order they appear herein, or may be executed in parallel. The sequence numbers are merely used to distinguish different operations, and the sequence numbers themselves do not represent any execution order. Additionally, these processes may include more or fewer operations, and these operations may be executed sequentially or in parallel. It should be noted that the descriptions such as "first," "second," etc., in this document are used to distinguish different messages, devices, modules, etc., and do not represent a sequential order, nor do they limit "first" and "second" to different types. "Multiple" means two or more, unless otherwise explicitly specified.
[0158] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods of the various embodiments of this application.
[0159] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein.
[0160] The above are merely preferred embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.
Claims
1. A remote procedure call (RPC) load balancing method, characterized by, include: In response to receiving an RPC request, the current allocation weight of each worker thread is obtained, wherein the allocation weight of each worker thread is dynamically adjusted based on the real-time load information of each worker thread and the predicted load information at future times; Worker threads are assigned to the RPC request based on the current allocation weight of each worker thread.
2. The method of claim 1, wherein, Also includes: At each first time interval, based on the real-time load information of each worker thread, predict the predicted load information of each worker thread at a future time. The allocation weight of each worker thread is adjusted based on the real-time load information of each worker thread and the predicted load information of each worker thread at future times.
3. The method of claim 2, wherein, The step of predicting the predicted load information of each worker thread at a future time based on the real-time load information of each worker thread includes: The real-time load information of each worker thread, the current time corresponding to the real-time load information, and the future time are input into the load prediction model for prediction, so as to obtain the predicted load information of each worker thread at the future time.
4. The method of claim 3, wherein, Also includes: At every second time interval, obtain the historical load information of each worker thread; The load prediction model is trained based on the historical load information of each worker thread.
5. The method of claim 4, wherein, The historical load information includes the actual load information of each worker thread at multiple times. Training the load prediction model based on the historical load information of each worker thread includes: The actual load information of each worker thread at the previous time, the previous time and the next time are input into the load prediction model for prediction, so as to obtain the predicted load information of each worker thread at the next time. The parameters of the load prediction model are adjusted based on the actual load information of each worker thread at the next time step and the predicted load information of each worker thread at the next time step.
6. The method of claim 2, wherein, The step of adjusting the allocation weight of each worker thread based on its real-time load information and its predicted load information at future times includes: Calculate the first weight of the worker thread based on the real-time load information and the predicted load information at a future time. The first weight of each worker thread is normalized to obtain the second weight of each worker thread; The allocation weight of each worker thread is adjusted according to the second weight of each worker thread.
7. The method of claim 1, wherein, Also includes: Real-time load information of each of the worker threads is collected and stored, along with the collection time. The real-time load information of the worker thread includes at least one of the following: CPU utilization, average response time, and task queue length.
8. The method according to any one of claims 1-7, characterized in that, After allocating worker threads to the RPC request based on the current allocation weight of each worker thread, the process further includes: The RPC request is processed by a worker thread allocated to it, and the processing result of the RPC request is obtained. Output the processing result of the RPC request.
9. The method of claim 8, wherein, The step of processing the RPC request by assigning a worker thread to the RPC request and obtaining the processing result of the RPC request includes: Add the RPC request to the task queue of the worker thread assigned to the RPC request; By assigning a worker thread to the RPC request, the RPC requests in the task queue are processed sequentially to obtain the processing result of the RPC request.
10. A method of RPC load balancing, characterized in that, Applied to a distributed storage server, the method includes: Receive RPC requests sent by client devices, wherein the RPC requests include at least one of the following: read data request, write data request, disk management request, snapshot management request, data migration request, and cluster management request; Obtain the current allocation weight of each worker thread, wherein the allocation weight of each worker thread is dynamically adjusted based on the real-time load information of each worker thread and the predicted load information at future times; Worker threads are assigned to the RPC request based on the current allocation weight of each worker thread.
11. The method of claim 10, wherein, Also includes: The RPC request is processed by a worker thread allocated to it, and the processing result of the RPC request is obtained. The processing result of the RPC request is returned to the client device.
12. The method according to claim 10 or 11, characterized in that, Also includes: At each first time interval, based on the real-time load information of each worker thread, predict the predicted load information of each worker thread at a future time. The allocation weight of each worker thread is adjusted based on the real-time load information of each worker thread and the predicted load information of each worker thread at future times.
13. A server, characterized by include: At least one processor; as well as A memory that is communicatively connected to the at least one processor; The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, cause the server to perform the method according to any one of claims 1-12.
14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, implement the method as described in any one of claims 1-12.
15. A computer program product comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1-12.