Load balancing method and device, electronic equipment and storage medium
By using a client-side load balancing method, the identification information of candidate server nodes is obtained, probe packets are sent and path quality is evaluated, and the target server node is selected. This solves the problems of low load balancing efficiency and robustness in distributed systems and achieves fast response and stable path selection.
Patent Information
- Application Number
- CN202512017106.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-30
- Publication Date
- 2026-01-30
- Estimated Expiration
- 2045-12-30
AI Technical Summary
In existing technologies, the load balancing efficiency and robustness of distributed systems are low, making it difficult to guarantee system stability. Especially in the context of highly dynamic and heterogeneous network environments, existing solutions are unable to respond quickly to network fluctuations or sudden traffic changes.
The load balancing method implemented on the client side obtains the identification information of candidate server nodes, sends probe packets and receives response messages, and combines the received timestamp and load rate information to determine the target transmission latency and path quality score, and selects the target server node.
It achieves stability and accuracy in path selection during network jitter or server load fluctuations, improves load balancing efficiency, responds quickly to network changes, reduces detection overhead, has scalability and robustness, and ensures the system's basic service capabilities.
Smart Images

Figure CN121441918A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of communication, and in particular to a load balancing method and device, an electronic device, and a storage medium. BACKGROUND
[0002] In cloud computing, edge computing, and large-scale distributed service architecture, the business requests of a client can usually be distributed to multiple server nodes for processing to achieve high throughput, low latency, and reliability guarantee. With the rapid development of the Internet of Things and mobile Internet, the network environment between the client and the server presents high dynamics and heterogeneity, such as network delay, jitter, packet loss rate, and significant fluctuations in server node load. Based on this, real-time sensing of network status and intelligent selection of optimal server nodes to improve the quality of service (QoS) become core problems.
[0003] Existing solutions mainly rely on centralized load decision at the center side, which is difficult to quickly respond to network fluctuations or sudden traffic changes, reduces the load balancing efficiency and robustness, and is difficult to guarantee the stability of the distributed system.
[0004] Therefore, how to improve the load balancing efficiency and robustness of the distributed system to improve the stability of the distributed system is one of the technical problems to be solved by the prior art. SUMMARY
[0005] To solve the problem of low load balancing efficiency and robustness of the distributed system in the prior art and difficulty in guaranteeing system stability, the embodiments of the present application provide a load balancing method, device, electronic device, and storage medium.
[0006] In a first aspect, the embodiments of the present application provide a load balancing method implemented at the client side, comprising: obtaining identification information of a candidate server node; sending a probe packet to the candidate server node according to the identification information of the candidate server node, the probe packet containing initiation timestamp information of the probe packet; receiving a response message returned by the candidate server node, the response message carrying a receiving timestamp and load rate information of the candidate server node; determining a target transmission delay between the client and the candidate server node according to the receiving timestamp, the initiation timestamp, and a historical transmission delay; determining a path quality score between the client and the candidate server node according to the load rate information of the candidate server node and the target transmission delay, the path quality score being used to represent the path quality between the client and the candidate server node; Based on the path quality score between the client and the candidate server node, a target server node is determined, and a service request is sent to the target server node.
[0007] Secondly, embodiments of this application provide a load balancing device implemented on the client side, comprising: The acquisition module is used to obtain the identification information of candidate server nodes; The sending module is used to send a probe packet to the candidate server node according to the identification information of the candidate server node, wherein the probe packet contains the initiation timestamp information of the probe packet; The first receiving module is used to receive the response message returned by the candidate server node, the response message carrying a receiving timestamp and the load rate information of the candidate server node; The first determining module is used to determine the target transmission delay between the client and the candidate server node based on the received timestamp, the initiation timestamp, and the historical transmission delay. The second determining module is used to determine the path quality score between the client and the candidate server node based on the load rate information of the candidate server node and the target transmission delay. The path quality score is used to characterize the path quality between the client and the candidate server node. The third determining module is used to determine the target server node based on the path quality score between the client and the candidate server node, so as to send a service request to the target server node.
[0008] Thirdly, embodiments of this application provide a load balancing method implemented on the global coordinator side, including: Receive a probe scheduling request sent by the client, wherein the probe scheduling request carries the requested service identification information; Based on the service identification information, candidate server nodes for providing corresponding services are determined; The client sends the identification information of the candidate server node to the client, enabling the client to send a probe packet to the candidate server node based on the identification information. The probe packet includes the initiation timestamp information of the probe packet. The client receives a response message returned by the candidate server node, which carries a reception timestamp and the load rate information of the candidate server node. Based on the reception timestamp, the initiation timestamp, and historical transmission delays, the client determines the target transmission delay between the client and the candidate server node. Based on the load rate information of the candidate server node and the target transmission delay, the client determines the path quality score between the client and the candidate server node. The path quality score is used to characterize the path quality between the client and the candidate server node. Based on the path quality score between the client and the candidate server node, the client determines the target server node and sends a service request to the target server node.
[0009] Fourthly, embodiments of this application provide a load balancing device implemented on the global coordinator side, comprising: The first receiving module is used to receive a probe scheduling request sent by the client, wherein the probe scheduling request carries the requested service identification information; The first determining module is used to determine the candidate server node that provides the corresponding service based on the service identification information; The sending module is configured to send the identification information of the candidate server node to the client, so that the client sends a probe packet to the candidate server node according to the identification information of the candidate server node, the probe packet containing the initiation timestamp information of the probe packet; receive a response message returned by the candidate server node, the response message carrying a reception timestamp and the load rate information of the candidate server node; determine the target transmission delay between the client and the candidate server node according to the reception timestamp, the initiation timestamp and historical transmission delay; determine the path quality score between the client and the candidate server node according to the load rate information of the candidate server node and the target transmission delay, the path quality score being used to characterize the path quality between the client and the candidate server node; and determine the target server node according to the path quality score between the client and the candidate server node, so as to send a service request to the target server node.
[0010] Fifthly, embodiments of this application provide an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the load balancing method described in this application.
[0011] Sixthly, embodiments of this application provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the load balancing method described in this application.
[0012] The beneficial effects of this application are as follows: The load balancing method, apparatus, electronic device, and storage medium provided in this application embodiment involve: a client obtaining the identification information of a candidate server node; sending a probe packet to the candidate server node based on the identification information, the probe packet containing the initiation timestamp information of the probe packet; receiving a response message returned by the candidate server node, the response message carrying the reception timestamp and the load rate information of the candidate server node; determining the target transmission delay between the client and the candidate server node based on the reception timestamp, the initiation timestamp, and historical transmission delay; determining the path quality score between the client and the candidate server node based on the load rate information of the candidate server node and the target transmission delay, the path quality score being used to characterize the path quality between the client and the candidate server node; and determining the target server node based on the path quality score between the client and the candidate server node, so as to send a service request to the target server node. In this embodiment, the client actively generates probe packets and sends them to candidate server nodes. Based on the probe packet's initiation timestamp, the candidate server node's returned reception timestamp, and historical transmission latency, the target transmission latency between the client and the candidate server node is determined. Path quality is assessed based on the real-time load rate returned by the candidate server node and the target transmission latency between the client and the candidate server node. This maintains the stability and accuracy of path selection even during network jitter or fluctuations in candidate server node load, thus enabling the client to proactively perceive the server node status in real-time and in multiple dimensions. This application transforms load balancing decision-making from a centralized decision-making process on the central side in existing technologies to a distributed intelligent decision-making process on the client side. The decision-making power for path quality score calculation and target server node selection is delegated to each client, forming a highly efficient distributed decision-making network. Clients can complete probe, path quality score calculation, and target server node selection locally within milliseconds, enabling rapid response to network jitter or sudden changes in server load, thus improving load balancing efficiency. Furthermore, the probe packet overhead is extremely low, avoiding new network congestion caused by introducing the probe mechanism. Each client's decision is independent and parallel; adding new clients or server nodes to the system will not put pressure on the system, demonstrating strong scalability. Furthermore, even if the global coordinator temporarily fails, each client can still perform effective load balancing based on local probes, ensuring the system's basic service capabilities and improving robustness.
[0013] Other features and advantages of this application will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the application. The objectives and other advantages of this application may be realized and obtained by means of the structures particularly pointed out in the written description, claims, and drawings. Attached Figure Description
[0014] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings: Figure 1 This is a schematic diagram illustrating an application scenario of the load balancing method provided in the embodiments of this application; Figure 2 A flowchart illustrating the load balancing method provided in an embodiment of this application; Figure 3 This application provides a schematic diagram of the process for setting the probe time offset for each client cluster in an embodiment of the present application. Figure 4 A schematic diagram illustrating the process of setting the detection frequency compression coefficient for each client cluster, provided for an embodiment of this application; Figure 5 A flowchart illustrating the process of determining the target transmission delay between the client and the candidate server node, provided in an embodiment of this application; Figure 6 A flowchart illustrating the load balancing method implemented on the client side as provided in the embodiments of this application; Figure 7 A schematic diagram of the structure of a load balancing device implemented on the client side as provided in an embodiment of this application; Figure 8 A flowchart illustrating the load balancing method implemented on the global coordinator side as provided in this application embodiment; Figure 9 A schematic diagram of the load balancing device implemented on the global coordinator side as provided in the embodiments of this application; Figure 10 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0015] To address the issues of low load balancing efficiency and robustness in existing distributed systems, which makes it difficult to guarantee system stability, this application provides a load balancing method, apparatus, electronic device, and storage medium.
[0016] The preferred embodiments of this application are described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit this application. Furthermore, the embodiments and features in the embodiments of this application can be combined with each other without conflict.
[0017] First refer to Figure 1 This is a schematic diagram of an application scenario of the load balancing method provided in the embodiments of this application. It may include a global coordinator 101, K client clusters 102 (cluster 1 to cluster K), and at least one server cluster 103. Each client cluster 102 includes multiple clients, and the server cluster 103 includes multiple server nodes. The K client clusters 102 are divided by the global coordinator 101 according to geographical regions, and each client cluster 102 is assigned a unique number, which can be represented by 1 to K.
[0018] The server node can be a standalone physical server or a cloud server that provides basic cloud computing services such as cloud servers, cloud databases, and cloud storage. The server, client, and global coordinator 101 are connected via a network; however, this embodiment does not limit the connection.
[0019] Based on the above application scenarios, the following will refer to the appendix. Figures 2-9 The exemplary embodiments of this application are described in more detail below. It should be noted that the above application scenarios are shown only to facilitate understanding of the spirit and principles of this application, and the implementation methods of this application are not limited in any way. On the contrary, the implementation methods of this application can be applied to any applicable scenario.
[0020] like Figure 2 The diagram shown illustrates the implementation flow of the load balancing method provided in this application embodiment. The load balancing method may include the following steps: S21. The client sends a probe scheduling request to the global coordinator, which carries the requested service identification information.
[0021] Initially, after dividing all managed clients into K clusters based on geographical regions and assigning a unique number to each client cluster, the global coordinator, to avoid network congestion caused by a large number of clients in multiple clusters simultaneously sending probe packets to the server node, and to improve probe efficiency and avoid resource waste, performs differentiated probe scheduling for different client clusters in both time and space dimensions. Different probe time offsets and probe frequency compression coefficients are assigned to each client cluster. Setting different probe time offsets for different client clusters allows clients in different clusters to probe at off-peak times, avoiding network congestion caused by synchronous probes. Setting different probe frequency compression coefficients for client clusters of different sizes compresses the probe frequency, allowing clients in client clusters of different sizes to send probe packets to the server node at different probe frequencies, improving probe efficiency and dynamically adapting to resource utilization.
[0022] When implementing, it can be done according to the following: Figure 3 The process shown involves setting the probe time offset for each client cluster, including the following steps: S31. The global coordinator determines the disturbance range based on the basic detection cycle.
[0023] In practice, the global coordinator can be set to the following disturbance range: ,in, The basic detection cycle can be set according to actual needs, such as 10 seconds. This application embodiment does not limit this.
[0024] S32. For each cluster, obtain a random perturbation value from the perturbation interval.
[0025] In specific implementation, for the first A cluster, the global coordinator can be generated from the disturbance interval Select a random perturbation value .
[0026] S33. Determine the detection time offset corresponding to the cluster based on the cluster number, basic detection period, and random disturbance value.
[0027] In practice, the global coordinator can calculate the first step using the following formula. The probe time offset corresponding to each client cluster:
[0028] in, Indicates the first The probe time offset corresponding to each cluster ; Indicates the first The cluster number, ; Indicates the basic detection period; Indicates the first Random perturbation values corresponding to each cluster , This represents the disturbance range.
[0029] In this way, the global coordinator can calculate the probe time offset for each client cluster.
[0030] For example, suppose there are 4 client clusters, numbered 1, 2, 3, and 4 respectively, with a basic detection cycle. If the time interval is seconds, then the disturbance range is... The random disturbance values are as follows: , , , The probe time offsets for each client cluster are as follows: , , , Therefore, the client in client cluster 1 can launch a probe packet at 9.5 seconds, the client in client cluster 2 can launch a probe packet at 20.3 seconds, the client in client cluster 3 can launch a probe packet at 29.2 seconds, and the client in client cluster 4 can launch a probe packet at 40.3 seconds.
[0031] S34. Send the detection time offset corresponding to each cluster to each client in the corresponding cluster.
[0032] In practice, the global coordinator distributes the probe time offset corresponding to each client cluster to each client in the corresponding cluster, and each client stores it locally.
[0033] When implementing, it can be done according to the following: Figure 4 The process shown describes how to set the probe frequency compression factor for each client cluster, including the following steps: S41, The global coordinator counts the number of clients in each cluster.
[0034] In practice, the global coordinator counts the number of clients in each client cluster.
[0035] S42. For each cluster, determine the corresponding detection frequency compression factor based on the number of clients in the cluster and the base cluster size.
[0036] In practice, the global coordinator can calculate the first step using the following formula. Compression factor of the probe frequency corresponding to each client cluster:
[0037] in, Indicates the first The compression factor of the detection frequency corresponding to each cluster; Indicates the first The number of clients contained in a cluster; Indicates the baseline cluster size.
[0038] The base cluster size, i.e. the number of clients included in the base cluster, can be set according to actual needs, such as 10,000 clients. This application embodiment does not limit this.
[0039] In this way, the global coordinator can determine the probe frequency compression factor for each client cluster.
[0040] For example, assuming a baseline cluster size , The number of clients in each client cluster is as follows (in seconds): (Small cluster) (Benchmark Cluster) (Medium-sized cluster) (Large cluster). The compression factor for the probe frequency corresponding to client cluster 1 is: Then its detection period is: Seconds (more frequent probes). The probe frequency compression factor for client cluster 2 is: Then its detection period is: Seconds (standard probe). The probe frequency compression factor for client cluster 3 is: Then its detection period is: Seconds (reducing the detection frequency). The detection frequency compression factor for client cluster 4 is: Then its detection period is: Seconds (lower detection frequency).
[0041] S43. Distribute the detection frequency compression coefficients corresponding to each cluster to each client in the corresponding cluster.
[0042] In practice, the global coordinator distributes the corresponding probe frequency compression coefficients for each cluster to each client in the corresponding cluster, and each client stores them locally.
[0043] In this way, the client can determine the detection time based on the detection time offset issued by the global coordinator, and compress the detection frequency proportionally according to the detection frequency compression factor.
[0044] Before a client needs to initiate a business request, it sends a probe scheduling request to the global coordinator. The probe scheduling request carries the business identification information of the request.
[0045] S22. The global coordinator determines the candidate server nodes that provide the corresponding services based on the service identification information.
[0046] In specific implementation, the global coordinator maintains a list of correspondences between different service identifiers and server nodes that can provide the corresponding services. When a probe scheduling request is received from a client, the global coordinator searches for server nodes that can provide the corresponding services from the list of correspondences between service identifiers and server nodes, and selects them as candidate server nodes. The candidate server nodes may include all server nodes corresponding to the service identifier. The global coordinator may also select multiple server nodes that are close to the client from the server nodes that can provide services for the service based on the client's geographical location. This application embodiment does not limit this.
[0047] S23. The global coordinator sends the identification information of the candidate server nodes to the client.
[0048] S24. The client sends a probe packet to the candidate server node based on the identification information of the candidate server node. The probe packet contains the timestamp information of the probe packet being initiated.
[0049] In practical implementation, after receiving the identification information of candidate server nodes sent by the global coordinator, the client generates a lightweight probe packet for each candidate server node before initiating a service request, based on the candidate server node's identification information, the probe packet's initiation timestamp, and the client's location hash value. The initiation timestamp represents the time when the probe packet is subsequently sent to the candidate server node. In this application, the lightweight probe packet is a network data packet constructed and sent during network communication for performing state detection tasks targeting server nodes; it is small in size, simple in structure, and has minimal impact on system load.
[0050] Specifically, the size of the lightweight probe packet can be set to no more than 10% of the standard service request, and it can carry the identifier and timestamp fields of the candidate server nodes, and be sent to each candidate server node in parallel through an encrypted channel.
[0051] In one implementation, the lightweight probe packet can adopt a layered encoding structure. The first layer is the packet header, used to store the identifier of the candidate server node, the timestamp of the probe packet initiation, and the protocol version number (such as TCP, UDP, etc.). The second layer is a variable-length payload field, used to store the client's location hash value and network access type (such as Wi-Fi, 5G, or wired network encoding). The third layer is a cyclic redundancy check (CRC) code, used for integrity verification at the application layer. The length of each layer can be set according to requirements. For example, the first layer can be set to 8 bytes, of which 2 bytes can be used to store the identifier of the candidate server node, 4 bytes can be used to store the timestamp accurate to microseconds, and 2 bytes are used to store the protocol version number; the length of the second layer can be set to no more than 40 bytes; the length of the third layer can be set to 4 bytes. This embodiment of the application does not limit this. (Probe packet length) It can be set to satisfy the following constraints: ,in, The probe packet represents the average length of historical business requests sent by the client to all server nodes. This ensures that the probe traffic is negligible compared to the corresponding business request traffic, preventing the probe packet from being excessively large even when calculated proportionally for large request scenarios (such as file uploads and video streaming). For example, 10% of a 1MB business request is 100KB. Due to the 52-byte constraint, the total length of the probe packet will not exceed 52 bytes, ensuring that the probe overhead is ≤10% and lower than the length of a regular control message. It should be noted that 52 bytes is just an example; it can be set according to requirements during implementation, and this application does not limit this. The structure of the probe packet can be designed with reference to the UDP minimum header (RFC768) and the Ethernet CRC-32 checksum specification (IEEE802.3-2018), combined with IPv4 / IPv6 MTU requirements for reduction (52 bytes is much smaller than the standard MTU (1500 bytes)).
[0052] After generating probe packets for each candidate server node, the client can determine the probe time based on the probe time offset issued by the global coordinator, and determine the probe frequency based on the probe frequency compression coefficient. Based on the probe time and probe frequency, the client can send the probe packets corresponding to each candidate server node to the candidate server node corresponding to the identifier of each candidate server node in parallel.
[0053] S25. After receiving the probe packet, the candidate server node returns a response message to the client. The response message carries the receiving timestamp and the load rate information of the candidate server node.
[0054] In one implementation, after receiving a probe packet from the client, each candidate server node generates a response message containing a reception timestamp and its load rate information, and returns it to the client. The reception timestamp is the time the probe packet was received. The load rate of the candidate server node can be obtained by weighted summation of current CPU utilization, memory utilization, and network bandwidth utilization; this embodiment does not limit this. The size of the response message can be limited, for example, to no more than 1.2 times the size of the probe packet; this embodiment does not limit this as well.
[0055] S26. The client determines the target transmission delay between the client and the candidate server node based on the received timestamp, the initiation timestamp, and the historical transmission delay.
[0056] In practice, for each candidate server node, the client extracts the receiving timestamp and the initiation timestamp from the response message returned by the candidate server node, and determines the target transmission delay between the client and the candidate server node based on the receiving timestamp, the initiation timestamp and the historical transmission delay.
[0057] When implementing, it can be done according to the following: Figure 5 The process shown determines the target transmission latency between the client and the candidate server node, including the following steps: S51. The client determines the transmission delay between the client and the candidate server node based on the received timestamp and the initiation timestamp.
[0058] In practice, for each candidate server node, the client subtracts the probe packet's initiation timestamp from the reception timestamp in the response message returned by the candidate server node to obtain the transmission delay between the client and that candidate server node.
[0059] S52. Obtain the historical transmission delay of the previously preset quantity.
[0060] In practice, the historical transmission delays preceding this one refer to the historical transmission delays corresponding to all probe packets sent by the client to all server nodes before initiating this probe packet. The preset number is M-1. Therefore, including the transmission delay of this one, there are a total of M consecutive transmission delays. The client records the transmission delay sequence of the M consecutive probes, and the Mth probe is the current probe.
[0061] S53. Obtain the weight of each historical transmission delay and the weight of the current transmission delay.
[0062] In practice, the client can calculate the first step using the following formula. Weights of each transmission delay:
[0063] in, Indicates the first One transmission delay, , This represents the total number of historical transmission delays and the current transmission delay.
[0064] By using the formula described above for calculating the weight of transmission delay, the historical transmission delay in the middle part can contribute more, effectively filtering out abnormal fluctuations while maintaining sensitivity to changes in delay trends, thus providing a stable delay input for the path quality scoring model.
[0065] In one implementation, the number of M can be adaptively adjusted according to the network jitter intensity, and the adjustment rule is as follows: The total number of historical transmission delays and the current transmission delay. The value of can be determined according to the following formula:
[0066] in, This represents the sensitivity coefficient, and V represents the current network jitter variance of the client. Used for control Sensitivity to the current network jitter variance V of the client.
[0067] The current network jitter variance V of the client can be calculated from the transmission delay sample window recorded by the client and determined using the exponentially weighted moving variance algorithm, reflecting the degree of current network delay fluctuation.
[0068] When network jitter is high (V is high), the window M is automatically increased to incorporate more historical data for filtering, making the latency estimation smoother and more stable. When the network is stable, the window M is decreased to make the evaluation more agile. The range of values can be set based on practical experience; for example, when When M is insensitive to network latency jitter, the window size M changes gradually. When the value of is large, such as when When M is constant, it means that the value of M is not sensitive to network latency jitter and the window size M changes gradually.
[0069] In practice, the value of M can also be preset according to the business type, and this application embodiment does not limit this.
[0070] S54. Based on the historical transmission delays and their weights, as well as the current transmission delay and its weight, calculate the weighted average to obtain the target transmission delay between the client and the candidate server node.
[0071] In practice, the client calculates the target transmission delay between the client and the candidate server node by weighting and averaging the historical transmission delays and their weights, as well as the current transmission delay and its weight.
[0072] To make the calculated target transmission delay more accurate, before calculating the weights of each historical transmission delay and the current transmission delay, the 3σ criterion can be used to remove outliers from these M transmission delays. Outliers that deviate from the mean by more than 3σ are removed, and the weighted average of the remaining transmission delays is taken as the effective transmission delay, i.e., the target transmission delay.
[0073] In this application, in order to cope with network latency jitter and sudden anomalies, a dynamic time window filtering algorithm is used to determine the value of M and the weight of each transmission delay when calculating the target transmission delay between the client and the candidate server node, and the weighted average value is calculated to obtain the effective transmission delay.
[0074] S27. The client determines the path quality score between the client and the candidate server nodes based on the load rate information of the candidate server nodes and the target transmission latency.
[0075] Among them, the path quality score is used to characterize the path quality between the client and the candidate server node.
[0076] In one implementation, the client can use the following path quality scoring model to calculate the client's relationship with the first... Path quality score between candidate server nodes:
[0077] in, Indicates the client and the first Path quality score between candidate server nodes , The number of candidate server nodes; Indicates the first Load rate of each candidate server node; Indicates the first The maximum load threshold corresponding to each candidate server node; Indicates the client and the first Target transmission latency between candidate server nodes; For the first Load rate of candidate server nodes Weighting coefficients; and For the client and the Target transmission delay between candidate server nodes The weighting coefficients.
[0078] The path quality assessment model is a crucial step in determining the final target server node selection. This model combines two core metrics: server node load rate and transmission latency, and uses weighting coefficients... , and A weighted fusion is performed to obtain a comprehensive path quality score.
[0079] In one implementation, there may be a problem of insufficient information in complex network environments. For example, when the available bandwidth, packet loss rate, or queue depth of a server node are highly correlated with service performance, failure to consider these indicators may lead to distorted evaluation results. Based on this, in order to improve the accuracy of path quality evaluation, after a candidate server node receives a probe packet sent by the client, it can return the candidate server node's packet loss rate information, available bandwidth information, and queue depth information along with the receiving timestamp and load rate information to the client.
[0080] Specifically, in addition to the receiving timestamp and the load rate information of the candidate server node, the response message may also include the candidate server's packet loss rate, available bandwidth, and queue depth. The queue depth is the number of pending service requests in the queue of service requests currently being processed by the candidate server node. The client can determine the path quality score between the client and the candidate server node based on the candidate server node's load rate, target transmission latency, packet loss rate, available bandwidth, and queue depth information.
[0081] At this point, the client can use the following path quality scoring model to calculate the client's performance relative to the first... Path quality score between candidate server nodes:
[0082] in, Indicates the client and the first Path quality score between candidate server nodes , The number of candidate server nodes; Indicates the first Load rate of each candidate server node; Indicates the first The maximum load threshold corresponding to each candidate server node; Indicates the client and the first Target transmission latency between candidate server nodes; Indicates the first Packet loss rate of each candidate server node; Indicates the first The available bandwidth of each candidate server node; Indicates the first The queue depth of each candidate server node; For the first Load rate of candidate server nodes Weighting coefficients; and For the client and the Target transmission delay between candidate server nodes Weighting coefficients; For the first Packet loss rate of candidate server nodes Weighting coefficients; For the first The weighting coefficient of the ratio of available bandwidth to queue depth for each candidate server node.
[0083] Weighting coefficients of packet loss rate The weighting factor for the ratio of available bandwidth to queue depth of candidate server nodes can be configured based on the business's sensitivity to network jitter. The path quality score is used to measure the ability of candidate server nodes to handle high-traffic services and can be set based on practical experience; this embodiment does not impose any limitations on this. Packet loss rate supplements the shortcomings of transmission latency in fully reflecting link stability. Joint modeling of available bandwidth and queue depth can reflect the adaptability of candidate server nodes to high-traffic services, and the introduction of a logarithmic function avoids the marginal effect when available bandwidth is too large. Through the fusion of multi-dimensional service quality indicators, the path quality score can more comprehensively reflect the service capabilities of candidate server nodes, providing a more reliable basis for service scheduling decisions.
[0084] In the two path quality scoring models mentioned above, the weighting coefficients , and The global coordinator can periodically adjust the performance based on historical business processing data of server nodes. To improve real-time adaptive capabilities, machine learning and optimization algorithms can be introduced. First, using a reinforcement learning framework, the client continuously tries different combinations of weight coefficients during network interactions, using throughput and transmission latency stability as reward signals to gradually learn the optimal weight allocation. Second, with the help of Bayesian optimization, the system quickly converges to the optimal solution in a continuous parameter space, thus avoiding the inefficiency of manual parameter tuning. Finally, the system can implement business-aware parameter configuration for different types of services: for example, improving performance in low-latency service scenarios. The weight value is increased in high-throughput scenarios. The weighting of the values is proportional to the bandwidth-related indicators, thereby enhancing the adaptability to complex network environments.
[0085] S28. The client determines the target server node based on the path quality score between the client and the candidate server nodes.
[0086] In practice, when the client determines that the difference between the highest path quality score and other path quality scores meets the anti-oscillation selection condition, it selects the candidate server node with the highest historical connection success rate from the candidate server nodes corresponding to the highest path quality score and the candidate server nodes corresponding to other path quality scores as the target server node.
[0087] Specifically, the anti-oscillation selection condition is determined to be met when the difference between the highest path quality score and other path quality scores meets the following condition:
[0088] in, Indicates the highest path quality score; This represents the quality score of any other path; Indicates the difference threshold. ,in, This represents the total number of times the client accesses (any) server node. Indicates the cumulative duration of network steady state. Indicates the steady-state reference duration of the network. This represents the adjustment constant. One or more paths can be rated as having the second-highest quality.
[0089] The client can determine whether the network is in a stable state by monitoring the short-term variance of transmission latency, and accumulate the duration of steady-state operation to obtain the cumulative steady-state duration. The network steady-state baseline duration is a preset constant used for standardization; for example, it can be set to, but is not limited to, 3600 seconds (1 hour). The adjustment constant k is used to control the difference threshold. The overall magnitude is to prevent it from being too large or too small, and its value range can be set according to needs, such as [1000, 10000]. This application embodiment does not limit this. When the client has rich experience in accessing ( Large) and the network is stable in the long term. When the difference threshold is large, The value increases, allowing for a wider range of path quality score differences, and tends to maintain existing choices, reducing unnecessary switching between candidate server nodes with similar scores. When client access frequency is low or the network is unstable, the difference threshold... A smaller value indicates that the client is more sensitive to differences in path quality scores and is more likely to switch to explore a better path.
[0090] When the oscillation selection condition is met, the client does not directly select the candidate server node with the highest path quality score, but instead prioritizes the candidate server node with a higher historical connection success rate. This can better balance historical access experience and network steady-state duration, thereby selecting a better path.
[0091] Considering future trends, one possible implementation could incorporate time series forecasting into this mechanism. For example, an ARIMA (Autoregressive Integrated Moving Average) model or an LSTM (Long Short-Term Memory network) model could be used to predict the transmission latency and load trends of candidate server nodes over a future period, thus mitigating potential short-term fluctuations. Simultaneously, a Markov chain model could be used to model node switching states, and the transition probability could be used to determine the presence of high-frequency switching risks. At the business level, the anti-oscillation threshold δ could be linked to SLA (Service Level Agreement) requirements. For instance, a smaller threshold could be used for real-time services to prioritize latency, while a larger threshold could be used for background batch processing services to reduce switching overhead. Through predictive enhancement and business-aware adjustment, the anti-oscillation mechanism can achieve a better balance between stability and flexibility.
[0092] S29. The client sends a service request to the target server node.
[0093] In practice, the client encapsulates the business request in a data frame carrying a path selection identifier (i.e., the identifier of the target server node) and sends it to the target server node.
[0094] After the client sends a service request to the target server node, it also includes: The client monitors the current transmission latency between itself and the target server node in real time. If the current transmission latency is determined to be greater than the latency threshold, the client determines the traffic splitting ratio based on the current transmission latency and the latency threshold. Subsequent business requests are then split to the candidate server nodes according to the splitting ratio. The candidate server nodes are the candidate server nodes with the highest path quality scores among the remaining candidate server nodes. The client also sends a link degradation alarm to the global coordinator.
[0095] In one implementation, the client can determine the latency threshold using the following formula:
[0096] in, Indicates the delay threshold; This represents the historical average transmission delay. This represents the standard deviation of historical transmission delay.
[0097] In one implementation, a delay threshold can be preset based on empirical values, but this application does not limit this approach.
[0098] The client can calculate the traffic splitting ratio using the following formula:
[0099] in, Indicates the diversion ratio; Indicates the shunting sensitivity factor; Indicates the current transmission delay; This indicates the time delay threshold.
[0100] Shunting Sensitive Factors It is an adjustable parameter greater than 0, which can be set according to the actual situation. This application embodiment does not limit it. The current transmission delay overtime threshold is the traffic splitting sensitivity factor. Used to control the split ratio For the current transmission delay timeout threshold The degree of sensitivity. The larger the value, the more sensitive the system is to latency degradation; only a small latency overscalar is needed. ), Diversion ratio It will quickly rise to a very high value, and vice versa. The smaller the value, the less sensitive the system is to latency degradation; even if the latency exceeds the scalar by a large amount, the offloading ratio increases slowly. The offloading ratio grows exponentially, ensuring a rapid response to any sudden increase in current transmission latency.
[0101] S210. After the business request is processed, the client obtains the business processing performance information of the target server node.
[0102] In practice, after a business request is processed, the client obtains its business processing performance information from the target server node, which may include: the target server node's transmission throughput information, packet loss rate information, and response latency information, etc.
[0103] S211. The client sends the business processing performance information of the target server node to the global coordinator.
[0104] In practice, the client sends the current business processing performance information of the target server node to the global coordinator.
[0105] S212. The global coordinator determines the health of the target server node based on the business processing performance information of the target server node, and updates the health of the target server node to the node health list.
[0106] In practical implementation, the global coordinator can determine the health of the target server node based on its current service processing performance information. Each service processing performance indicator can be judged using a threshold method, or a comprehensive health score can be obtained by weighted summing the ratios of the actual values of each service processing performance indicator to the normal baseline values. This embodiment does not limit this approach. The normal baseline values for service processing performance indicators can be set according to requirements, such as the difference between the highest and lowest values within the normal range of the indicator. Subsequently, the health of the target server node is updated in the node health list.
[0107] S213. The global coordinator will broadcast the updated node health list to clients in each cluster during the next list update cycle.
[0108] In practice, the inventory update cycle must meet the following conditions:
[0109] Where T represents the inventory update cycle; Indicates the frequency of network fluctuations; Indicates the number of online clients; Indicates the baseline client size.
[0110] The aforementioned inventory update cycle adopts a logarithmic function, which ensures that the inventory update cycle grows slowly and controllably as the client scale increases.
[0111] In this way, by viewing the node health list, clients in each cluster can prioritize candidate server nodes with high health when sending probe packets to candidate server nodes, thereby improving load balancing efficiency.
[0112] After the business request is processed, the client obtains the business processing performance information of the target server node and sends it to the global coordinator. The global coordinator can then update the weighting coefficients based on the business processing performance information. , and and update the weight coefficients , and Send to the client, which uses the updated weighting coefficients. , and Update the path quality scoring model.
[0113] In this embodiment, the global coordinator is responsible for the detection scheduling of the client cluster and the health management of the server nodes. It allocates detection time offset and detection frequency compression coefficient to each cluster, broadcasts the server node health list, periodically generates the server node health list, and updates the period to ensure balanced global scheduling when the number of clients changes or the network fluctuates.
[0114] In the load balancing method provided in this application, the client actively generates a lightweight probe packet and sends it to the candidate server node. Based on the probe packet's initiation timestamp, the candidate server node's return reception timestamp, and historical transmission latency, the target transmission latency between the client and the candidate server node is determined. Path quality is evaluated based on the real-time load rate returned by the candidate server node and the target transmission latency between the client and the candidate server node. This maintains the stability and accuracy of path selection even during network jitter or fluctuations in the candidate server node's load, thereby enabling the client to proactively perceive the server node's status in real time and in multiple dimensions. This application transforms load balancing decision-making from a centralized decision-making process on the central side in existing technologies to a distributed intelligent decision-making process on the client side. It decentralizes the decision-making power for path quality score calculation and target server node selection to each client, forming a highly efficient distributed decision-making network. The client can complete probing, path quality score calculation, and target server node selection locally within milliseconds, enabling rapid response to network jitter or sudden changes in server load, thus improving load balancing efficiency. Furthermore, the lightweight probe packet has extremely low overhead, avoiding new network congestion caused by introducing the probe mechanism. Each client makes independent and parallel decisions. Adding new client or server nodes to the system will not put pressure on it, demonstrating strong scalability. Furthermore, even if the global coordinator temporarily fails, each client can still perform effective load balancing based on local probing, ensuring the system's basic service capabilities and improving robustness.
[0115] Based on the same inventive concept, this application also provides a load balancing method implemented on the client side. Since the principle of the load balancing method implemented on the client side is similar to that of the load balancing method, the implementation of the load balancing method implemented on the client side can refer to the implementation of the load balancing method, and the repeated parts will not be described again.
[0116] like Figure 6 The diagram shown is a flowchart illustrating a client-side load balancing method implemented according to an embodiment of this application, which may include the following steps: S61. The client obtains the identification information of the candidate server node.
[0117] S62. Send a probe packet to the candidate server node according to the identification information of the candidate server node. The probe packet contains the timestamp information of the probe packet initiation. S63. Receive the response message returned by the candidate server node, the response message carrying the receiving timestamp and the load rate information of the candidate server node.
[0118] S64. Determine the target transmission delay between the client and the candidate server node based on the received timestamp, the initiation timestamp, and the historical transmission delay.
[0119] S65. Based on the load rate information of the candidate server nodes and the target transmission latency, determine the path quality score between the client and the candidate server nodes.
[0120] Among them, the path quality score is used to characterize the path quality between the client and the candidate server node.
[0121] S66. Based on the path quality score between the client and the candidate server nodes, determine the target server node and send the business request to the target server node.
[0122] In one implementation, determining the target transmission delay between the client and the candidate server node based on the received timestamp, the initiation timestamp, and the historical transmission delay specifically includes: Based on the received timestamp and the initiated timestamp, the transmission delay between the client and the candidate server node for this transaction is determined. Get the historical transmission latency of the previously preset quantity; Obtain the weight of each historical transmission delay and the weight of the current transmission delay; The target transmission delay between the client and the candidate server node is obtained by weighting and averaging the historical transmission delays and their weights, along with the current transmission delay and its weight.
[0123] In one implementation, the response message also carries packet loss rate information, available bandwidth information, and queue depth information of the candidate server; then
[0124] Based on the load rate information of the candidate server nodes and the target transmission latency, a path quality score is determined between the client and the candidate server nodes, specifically including: Based on the load rate information of the candidate server node, the target transmission latency, the packet loss rate information of the candidate server, the available bandwidth information, and the queue depth information, the path quality score between the client and the candidate server node is determined.
[0125] In one implementation, the target server node is determined based on the path quality score between the client and the candidate server node, specifically including: When the difference between the highest path quality score and other path quality scores is determined to meet the anti-oscillation selection condition, the candidate server node with the highest historical connection success rate is selected as the target server node from the candidate server node corresponding to the highest path quality score and the candidate server nodes corresponding to the other path quality scores.
[0126] In one implementation, after sending the service request to the target server node, the method further includes: Real-time detection of the current transmission latency between the target server node and the target server node; If it is determined that the current transmission delay is greater than the delay threshold, then the traffic splitting ratio is determined based on the current transmission delay and the delay threshold; Subsequent business requests will be routed to candidate server nodes according to the aforementioned traffic distribution ratio. These candidate server nodes are the ones with the highest path quality scores among the remaining candidate server nodes. Send a link degradation alarm message to the global coordinator.
[0127] In one implementation, before obtaining the identification information of the candidate server node, the following steps are also included: Receive the probe time offset and probe frequency compression factor sent by the global coordinator; The detection time is determined based on the detection time offset, and the detection frequency is determined based on the detection frequency compression factor; and Sending probe packets to the candidate server nodes based on their identification information, specifically including: The probe packet is sent to the candidate server node corresponding to the identifier of the candidate server node according to the probe time and the probe frequency.
[0128] Based on the same inventive concept, this application also provides a load balancing device implemented on the client side. Since the principle of the load balancing device implemented on the client side is similar to that of the load balancing method, the implementation of the load balancing device implemented on the client side can refer to the implementation of the load balancing method. Repeated parts will not be described again.
[0129] like Figure 7 As shown, it is a structural schematic diagram of a load balancing device implemented on the client side according to an embodiment of this application, which may include: The acquisition module 71 is used to acquire the identification information of the candidate server nodes; The sending module 72 is used to send a probe packet to the candidate server node according to the identification information of the candidate server node, wherein the probe packet contains the initiation timestamp information of the probe packet; The first receiving module 73 is used to receive a response message returned by the candidate server node, the response message carrying a receiving timestamp and the load rate information of the candidate server node; The first determining module 74 is used to determine the target transmission delay between the client and the candidate server node based on the receiving timestamp, the initiation timestamp, and the historical transmission delay. The second determining module 75 is used to determine the path quality score between the client and the candidate server node based on the load rate information of the candidate server node and the target transmission delay. The path quality score is used to characterize the path quality between the client and the candidate server node. The third determining module 76 is used to determine the target server node based on the path quality score between the client and the candidate server node, so as to send a service request to the target server node.
[0130] In one implementation, the first determining module 74 is specifically configured to: determine the current transmission delay between the client and the candidate server node based on the receiving timestamp and the initiation timestamp; obtain a preset number of historical transmission delays prior to this transmission delay; obtain the weight of each historical transmission delay and the weight of the current transmission delay; and calculate a weighted average based on the historical transmission delays and their weights, as well as the current transmission delay and its weights, to obtain the target transmission delay between the client and the candidate server node.
[0131] In one implementation, the response message also carries packet loss rate information, available bandwidth information, and queue depth information of the candidate server; then
[0132] The second determining module 75 is specifically used to determine the path quality score between the client and the candidate server node based on the load rate information of the candidate server node, the target transmission delay, the packet loss rate information of the candidate server, the available bandwidth information, and the queue depth information.
[0133] In one implementation, the third determining module 76 is specifically used to select the candidate server node with the highest historical connection success rate from the candidate server node corresponding to the highest path quality score and the candidate server nodes corresponding to the other path quality scores when the difference between the highest path quality score and other path quality scores satisfies the anti-oscillation selection condition.
[0134] In one implementation, the apparatus further includes: The detection module is used to detect the current transmission latency between the target server node and the target server node in real time after sending a service request to the target server node. The fourth determining module is used to determine the traffic splitting ratio based on the current transmission delay and the delay threshold if it is determined that the current transmission delay is greater than the delay threshold. The traffic splitting module is used to split subsequent business requests to alternative server nodes according to the splitting ratio. The alternative server nodes are the candidate server nodes with the highest path quality scores among the remaining candidate server nodes. The alarm module is used to send link degradation alarm information to the global coordinator.
[0135] In one implementation, the apparatus further includes: The second receiving module is used to receive the probe time offset and probe frequency compression coefficient sent by the global coordinator before obtaining the identification information of the candidate server node. The fifth determining module is used to determine the detection time based on the detection time offset and to determine the detection frequency based on the detection frequency compression coefficient; and The sending module is specifically used to send the probe packet to the candidate server node corresponding to the identifier of the candidate server node according to the probe time and the probe frequency.
[0136] Based on the same inventive concept, this application also provides a load balancing method implemented on the global coordinator side. Since the principle of solving the problem by the load balancing method implemented on the global coordinator side is similar to that of the load balancing method described above, the implementation of the load balancing method implemented on the global coordinator side can refer to the implementation of the load balancing method described above, and the repeated parts will not be described again.
[0137] like Figure 8 The diagram shown is a flowchart illustrating a load balancing method implemented on the global coordinator side according to an embodiment of this application, which may include the following steps: S81. The global coordinator receives a probe scheduling request sent by the client. The probe scheduling request carries the requested service identification information.
[0138] S82. Determine the candidate server nodes that provide the corresponding services based on the business identification information.
[0139] S83. Send the identification information of the candidate server node to the client, so that the client can send a probe packet to the candidate server node according to the identification information of the candidate server node. The probe packet contains the initiation timestamp information of the probe packet. Receive the response message returned by the candidate server node. The response message carries the reception timestamp and the load rate information of the candidate server node. Determine the target transmission delay between the client and the candidate server node according to the reception timestamp, initiation timestamp, and historical transmission delay. Determine the path quality score between the client and the candidate server node according to the load rate information of the candidate server node and the target transmission delay. The path quality score is used to characterize the path quality between the client and the candidate server node. Determine the target server node according to the path quality score between the client and the candidate server node, and send a service request to the target server node.
[0140] In one implementation, before receiving the probe scheduling request sent by the client, the following is also included: All managed clients are divided into K clusters according to geographical regions, and each cluster is assigned a unique number; The disturbance range is determined based on the basic detection period; For each cluster, a random perturbation value is obtained from the perturbation interval; The detection time offset corresponding to the cluster is determined based on the cluster number, the basic detection period, and the random disturbance value. The probe time offset corresponding to each cluster is sent to each client in the corresponding cluster.
[0141] In one implementation, the method further includes: Count the number of clients in each cluster; For each cluster, the detection frequency compression factor corresponding to the cluster is determined based on the number of clients contained in the cluster and the base cluster size. The compression coefficient of the detection frequency corresponding to each cluster is sent to each client in the corresponding cluster.
[0142] In one implementation, the method further includes: The system receives the service processing performance information of the target server node sent by the client, wherein the service processing performance information of the target server node is obtained by the client after the service request is processed. The health of the target server node is determined based on the service processing performance information of the target server node. Update the health status of the target server node to the node health list; In the next list update cycle, the updated node health list will be broadcast to clients in each cluster.
[0143] Based on the same inventive concept, this application also provides a load balancing device implemented on the global coordinator side. Since the principle of the load balancing device implemented on the global coordinator side is similar to that of the load balancing method, the implementation of the load balancing device implemented on the global coordinator side can refer to the implementation of the load balancing method. Repeated parts will not be described again.
[0144] like Figure 9 As shown, it is a structural schematic diagram of a load balancing device implemented on the global coordinator side according to an embodiment of this application, which may include: The first receiving module 91 is used to receive a probe scheduling request sent by the client, wherein the probe scheduling request carries the requested service identification information; The first determining module 92 is used to determine the candidate server node that provides the corresponding service based on the service identification information; The sending module 93 is configured to send the identification information of the candidate server node to the client, so that the client sends a probe packet to the candidate server node according to the identification information of the candidate server node, the probe packet containing the initiation timestamp information of the probe packet; receive a response message returned by the candidate server node, the response message carrying a reception timestamp and the load rate information of the candidate server node; determine the target transmission delay between the client and the candidate server node according to the reception timestamp, the initiation timestamp and the historical transmission delay; determine the path quality score between the client and the candidate server node according to the load rate information of the candidate server node and the target transmission delay, the path quality score being used to characterize the path quality between the client and the candidate server node; and determine the target server node according to the path quality score between the client and the candidate server node, so as to send a service request to the target server node.
[0145] In one implementation, the apparatus further includes: The partitioning module is used to divide all managed clients into K clusters according to geographical regions before receiving the probe scheduling request sent by the client, and assign a unique number to each cluster. The second determination module is used to determine the disturbance range based on the basic detection period; The acquisition module is used to acquire a random perturbation value from the perturbation interval for each cluster; The third determining module is used to determine the detection time offset corresponding to the cluster based on the cluster number, the basic detection period, and the random disturbance value. The first distribution module is used to distribute the probe time offset corresponding to each cluster to each client in the corresponding cluster.
[0146] In one implementation, the method further includes: The statistics module is used to count the number of clients in each cluster; The fourth determining module is used to determine the detection frequency compression coefficient corresponding to each cluster based on the number of clients contained in the cluster and the base cluster size. The second distribution module is used to distribute the detection frequency compression coefficients corresponding to each cluster to each client in the corresponding cluster.
[0147] In one implementation, the apparatus further includes: The second receiving module is used to receive the service processing performance information of the target server node sent by the client. The service processing performance information of the target server node is obtained by the client after the service request is processed. The fifth determining module is used to determine the health of the target server node based on the service processing performance information of the target server node; The update module is used to update the health status of the target server node to the node health status list; The broadcast module is used to broadcast the updated node health list to clients in each cluster during the next list update cycle.
[0148] Based on the same technical concept, this application also provides an electronic device 1000, referring to... Figure 10 As shown, the electronic device 1000 is used to implement the load balancing method described in the above-described method embodiments. The electronic device 1000 in this embodiment may include: a memory 1001, a processor 1002, and a computer program, such as a load balancing program, stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps in the various load balancing method embodiments described above.
[0149] This application embodiment does not limit the specific connection medium between the memory 1001 and the processor 1002. This application embodiment... Figure 10 The memory 1001 and the processor 1002 are connected via a bus 1003, and the bus 1003 is in Figure 10 The connections between other components are shown in bold and are for illustrative purposes only, not as limiting information. The bus 1003 can be divided into address bus, data bus, control bus, etc. For ease of illustration, Figure 10 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.
[0150] Memory 1001 may be volatile memory, such as random-access memory (RAM); memory 1001 may also be non-volatile memory, such as read-only memory, flash memory, hard disk drive (HDD), or solid-state drive (SSD); or memory 1001 may be any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but is not limited thereto. Memory 1001 may be a combination of the above-mentioned memories.
[0151] The processor 1002 is used to implement the load balancing method provided in the embodiments of this application.
[0152] This application also provides a computer-readable storage medium storing computer-executable instructions required to execute the processor, including a program required to execute the processor.
[0153] In some possible implementations, various aspects of the load balancing method provided in this application may also be implemented as a program product comprising program code that, when the program product is run on an electronic device, causes the electronic device to perform the steps in the load balancing method according to the various exemplary embodiments of this application described above.
[0154] Those skilled in the art will understand that embodiments of this application can be provided as methods, apparatus, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0155] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (devices), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the process. Figure 1One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0156] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0157] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0158] Although preferred embodiments of this application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of this application.
[0159] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.
Claims
1. A load balancing method, characterized by, The method applied to a client, comprising: obtaining identification information of a candidate server node; sending a probe packet to the candidate server node according to the identification information of the candidate server node, the probe packet containing an initiation time stamp information of the probe packet; receiving a response message returned by the candidate server node, the response message carrying a receiving time stamp and load rate information of the candidate server node; determining a target transmission time delay between the client and the candidate server node according to the receiving time stamp, the initiation time stamp and historical transmission time delays; determining a path quality score between the client and the candidate server node according to the load rate information of the candidate server node and the target transmission time delay, the path quality score being used to represent a path quality between the client and the candidate server node; determining a target server node according to the path quality score between the client and the candidate server node, to send a service request to the target server node.
2. The method of claim 1, wherein, The method further comprises: determining a current transmission time delay between the client and the candidate server node according to the receiving time stamp and the initiation time stamp; obtaining a preset number of historical transmission time delays before the current transmission time delay; obtaining a weight of each historical transmission time delay and a weight of the current transmission time delay; performing weighted average according to the historical transmission time delays and their weights and the current transmission time delay and its weight, to obtain the target transmission time delay between the client and the candidate server node.
3. The method of claim 1, wherein, The response message further carries packet loss rate information, available bandwidth information and queue depth information of the candidate server; and The method further comprises: determining the path quality score between the client and the candidate server node according to the load rate information of the candidate server node, the target transmission time delay, the packet loss rate information, the available bandwidth information and the queue depth information of the candidate server.
4. The method of claim 1, wherein, The method further comprises: when a difference between the highest path quality score and other path quality scores meets an anti-oscillation selection condition, selecting a candidate server node with the highest historical connection success rate from the candidate server node corresponding to the highest path quality score and the candidate server nodes corresponding to the other path quality scores as the target server node.
5. The method of claim 1, wherein, The method further comprises: after sending the service request to the target server node, real-time detecting a current transmission time delay between the client and the target server node; if the current transmission time delay is greater than a time delay threshold, determining a shunting ratio according to the current transmission time delay and the time delay threshold. shunt subsequent service requests to the alternative server node according to the shunt ratio, the alternative server node being the candidate server node with the highest path quality score among the remaining candidate server nodes; and sending link degradation alarm information to the global coordinator.
6. The method of claim 1, wherein, Before obtaining the identification information of the candidate server node, further comprising: receiving the probe time offset and the probe frequency compression coefficient issued by the global coordinator; determining the probe time according to the probe time offset and determining the probe frequency according to the probe frequency compression coefficient; and sending the probe packet to the candidate server node according to the identification information of the candidate server node, specifically comprising: sending the probe packet to the candidate server node corresponding to the identification information of the candidate server node according to the probe time and the probe frequency.
7. A load balancing method characterized by, The method applied to the global coordinator, comprising: receiving the probe scheduling request sent by the client, the probe scheduling request carrying the requested service identification information; determining the candidate server node providing the corresponding service according to the service identification information; sending the identification information of the candidate server node to the client, so that the client sends a probe packet to the candidate server node according to the identification information of the candidate server node, the probe packet containing the initiation timestamp information of the probe packet; receiving the response message returned by the candidate server node, the response message carrying the receiving timestamp and the load rate information of the candidate server node; determining the target transmission time delay between the client and the candidate server node according to the receiving timestamp, the initiation timestamp and the historical transmission time delay; determining the path quality score between the client and the candidate server node according to the load rate information of the candidate server node and the target transmission time delay, the path quality score representing the path quality between the client and the candidate server node; determining the target server node according to the path quality score between the client and the candidate server node, so as to send a service request to the target server node.
8. The method of claim 7, wherein, Before receiving the probe scheduling request sent by the client, further comprising: dividing all managed clients into K clusters according to geographical areas, and assigning a unique number to each cluster; determining a perturbation interval based on a basic probe period; for each cluster, obtaining a random perturbation value from the perturbation interval; determining the probe time offset corresponding to the cluster according to the number of the cluster, the basic probe period and the random perturbation value; issuing the probe time offset corresponding to each cluster to each client in the corresponding cluster.
9. The method of claim 8, wherein, Further comprising: counting the number of clients included in each cluster; for each cluster, determining the probe frequency compression coefficient corresponding to the cluster according to the number of clients included in the cluster and the reference cluster size; issuing the probe frequency compression coefficient corresponding to each cluster to each client in the corresponding cluster.
10. The method of claim 8, wherein, Further comprising: receive service processing performance information of the target server node sent by the client, the service processing performance information of the target server node being obtained by the client after completion of the service request processing; determine the health degree of the target server node based on the service processing performance information of the target server node; update the health degree of the target server node to a node health degree list; broadcast the updated node health degree list to the clients in each cluster in the next list update period.
11. A load balancing apparatus, characterized by, The application is applied to a client, and the device comprises: an obtaining module, configured to obtain identification information of a candidate server node; a sending module, configured to send a probe packet to the candidate server node according to the identification information of the candidate server node, the probe packet containing initiation timestamp information of the probe packet; a first receiving module, configured to receive a response message returned by the candidate server node, the response message carrying reception timestamp and load rate information of the candidate server node; a first determining module, configured to determine target transmission time delay between the client and the candidate server node according to the reception timestamp, the initiation timestamp and historical transmission time delay; a second determining module, configured to determine a path quality score between the client and the candidate server node according to the load rate information of the candidate server node and the target transmission time delay, the path quality score being used to represent path quality between the client and the candidate server node; a third determining module, configured to determine a target server node according to the path quality score between the client and the candidate server node, so as to send a service request to the target server node.
12. A load balancing apparatus, characterized by, The application is applied to a global coordinator, and the device comprises: a first receiving module, configured to receive a probe scheduling request sent by a client, the probe scheduling request carrying requested service identification information; a first determining module, configured to determine a candidate server node providing corresponding service according to the service identification information; a sending module, configured to send identification information of the candidate server node to the client, so that the client sends a probe packet to the candidate server node according to the identification information of the candidate server node, the probe packet containing initiation timestamp information of the probe packet; receive a response message returned by the candidate server node, the response message carrying reception timestamp and load rate information of the candidate server node; determine target transmission time delay between the client and the candidate server node according to the reception timestamp, the initiation timestamp and historical transmission time delay; determine a path quality score between the client and the candidate server node according to the load rate information of the candidate server node and the target transmission time delay, the path quality score being used to represent path quality between the client and the candidate server node; determine a target server node according to the path quality score between the client and the candidate server node, so as to send a service request to the target server node.
13. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor implements the load balancing method as claimed in any one of claims 1-10 when executing the program.
14. A computer-readable storage medium having stored thereon a computer program, characterized in that, The program is executed by the processor to implement the steps in the load balancing method as claimed in any one of claims 1-10.
Citation Information
Patent Citations
Business server side load balancing method, client side, server side and system
CN104079630A
Node display method and device, storage medium and program product
CN114490658A
Micro-service dynamic adaptive client load balancing method and system
CN117155942A
Proxy gateway scheduling method and device, storage medium and computer equipment
CN120711016A
Business scheduling method and device, computer equipment, readable storage medium and program product
CN120916198A