Method for realizing load balancing, load balancing server and cluster system

By adding request allocation weights to the computing server and combining historical request data for load balancing judgment and prediction, the problem of unbalanced load of the computing server is solved, reasonable allocation of user requests and session continuity is achieved, and user experience and business logic coherence is improved.

CN120234147APending Publication Date: 2025-07-01SHANGHAI ANHEHAO INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510362377.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

In the prior art, the computing server has the problem of load imbalance when processing user requests, and the lack of reasonable user request planning and allocation, resulting in some servers being overloaded while other servers being insufficient, and it is impossible to effectively predict and allocate the user request volume.

Method used

By adding request allocation weights to the calculation server based on the parameter data of the calculation server and the weighted polling strategy, and combining the request data within multiple historical calculation cycles, the request load balancing between the calculation servers is judged. If uneven, predict the amount of user requests in the current calculation cycle, and determine its request allocation amount in the current cycle based on the request allocation weight of the calculation server, and finally push the user request to the matching calculation server.

Benefits of technology

It realizes balancing judgment and allocation of requested loads between computing servers, ensures load balancing of computing servers in the current cycle, and maintains the continuity of user request sessions by analyzing user request stability and predicting the number of requests, and improves the logical coherence of user experience and business requests.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120234147A_ABST
    Figure CN120234147A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of computer clusters, and provides a method for realizing load balancing, a load balancing server and a cluster system.The method comprises the steps that a request distribution weight is added to a computing server; determining an ideal request amount of each computing server in a plurality of historical computing periods, comparing and analyzing the ideal request amount with an actual request amount, and judging whether request loads among the computing servers are balanced or not; predicting the user request amount in the current calculation period, and determining the request allocation amount of each calculation server in the current calculation period in combination with the request allocation weight of the calculation server; determining a predicted request quantity of the user in the current calculation period according to the request quantity of the user sending the request to the calculation server in the plurality of historical calculation periods; and according to the predicted request amount of the user and the residual request allocation amount of each computing server, pushing the user request to the matched computing server in the current computing period, thereby realizing load balancing management of the computing servers.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer clusters, and specifically relates to a method for implementing load balancing, a load balancing server, and a cluster system. Background Art

[0002] As one of the important modern technical systems, computing servers are crucial for the efficient, secure, and elastic processing of user requests, and their importance runs through all aspects of Internet services. However, there are still many deficiencies in the operation and management of computing servers. For example, problems such as load imbalance and insufficient session continuity exist. Therefore, it is of great significance to study a method for implementing load balancing, a load balancing server, and a cluster system.

[0003] In the prior art, when processing user requests, computing servers often suffer from problems such as uneven user requests and lack of reasonable user request planning and allocation. For example, some computing servers receive and process a large number of user requests, exceeding the load of the computing servers, while some computing servers receive and process a small number of user requests, far lower than the load of the computing servers. There is a lack of reasonable planning and allocation of user requests based on the user request volume and the performance of the computing servers, and there is also a lack of prediction of the user request volume within the computing cycle and the request volume of individual users based on the number of user requests and the stability of user requests within the historical computing cycle, so as to achieve reasonable allocation of user requests among computing servers within the computing cycle while maintaining the persistence of user request sessions, improving the user experience, and enhancing the logical coherence of business requests.

[0004] Therefore, the present invention provides a method for implementing load balancing, a load balancing server, and a cluster system. Summary of the Invention

[0005] In order to make up for the deficiencies of the prior art and solve at least one of the technical problems proposed in the background art.

[0006] The technical solution adopted by the present invention to solve its technical problems is as follows:

[0007] In a first aspect, the present invention provides a method for implementing load balancing, including:

[0008] S1: Adding request allocation weights to computing servers according to the parameter data of the computing servers and in combination with the weighted round-robin strategy;

[0009] S2: Determining the ideal request amount of each computing server in multiple historical computing cycles according to the request allocation weights, comparing and analyzing it with the actual request amount, and judging whether the request loads among the computing servers are balanced;

[0010] S3: If it is unbalanced, predict the user request volume within the current calculation cycle, and determine the request allocation amount for each computing server within the current calculation cycle in combination with the request allocation weight of the computing server;

[0011] S4: Analyze and judge the stability of user request sending based on the request volume of the user sending requests to the computing server in multiple historical calculation cycles, and determine the predicted request volume of the user within the current calculation cycle;

[0012] S5: Push the user request to the matching computing service within the current calculation cycle according to the predicted request volume of the user and the remaining request allocation amount of each computing server.

[0013] In a second aspect, the present invention provides a load balancing server, including:

[0014] Server weight calculation module: Add a request allocation weight for the computing server according to the parameter data of the computing server and in combination with the weighted round-robin strategy;

[0015] Load balancing judgment module: Determine the ideal request amount of each computing server in multiple historical calculation cycles according to the request allocation weight, and compare and analyze it with the actual request amount to judge whether the request load among the computing servers is balanced;

[0016] User request allocation module: If it is unbalanced, predict the user request volume within the current calculation cycle, and determine the request allocation amount for each computing server within the current calculation cycle in combination with the request allocation weight of the computing server;

[0017] User request prediction module: Analyze and judge the stability of user request sending based on the request volume of the user sending requests to the computing server in multiple historical calculation cycles, and determine the predicted request volume of the user within the current calculation cycle;

[0018] User request push module: Push the user request to the matching computing server within the current calculation cycle according to the predicted request volume of the user and the remaining request allocation amount of each computing server.

[0019] In a third aspect, the present invention provides a load balancing server, including:

[0020] Load balancing server, computing server and user;

[0021] The computing server is used to receive user requests;

[0022] The load balancing server is used to assign weights to requests for computing servers, determine whether the request loads among the computing servers are balanced, and in the case of imbalance, determine the request allocation amount for each computing server in the current computing cycle and predict the predicted request volume of users in the current computing cycle, so as to implement user request pushing.

[0023] The beneficial effects of the present invention are as follows:

[0024] 1. According to the parameter data of the computing servers, determine the request allocation weights of the computing servers, and based on the request allocation weights, determine the ideal request amounts of each computing server in multiple historical computing cycles, and compare and analyze them with the actual request amounts to determine whether the request loads among the computing servers are balanced. If they are not balanced, predict the user request volume in the current computing cycle, and in combination with the request allocation weights of the computing servers, determine the request allocation amount for each computing server in the current computing cycle. The present invention realizes the judgment of whether the request loads among the computing servers are balanced, and in the case of imbalance, determines the request allocation amount of the computing servers in combination with the request allocation weights of the computing servers and the predicted user request volume in the current computing cycle, ensuring the load balance of the computing servers in the current computing cycle.

[0025] 2. According to the request volumes of the users sending requests to the computing servers in multiple historical computing cycles, analyze and judge the stability of the user request sending, and determine the predicted request volume of the users in the current computing cycle. According to the predicted request volume of the users and the remaining request allocation amounts of each computing server, in the current computing cycle, push the users to the matching computing servers. The present invention realizes the matching push of user requests by analyzing the stability of user requests, predicting the request volume of users, and combining the remaining request amounts of the computing servers, further realizing the load balance of the requests of the computing servers. And pushing the users to the matching computing servers according to the predicted request volume of the users can maintain the continuity of the user request session, improve the user experience and the logical coherence of business requests. Description of the Drawings

[0026] The present invention will be further described below with reference to the drawings.

[0027] Figure 1 is the step flowchart of a method for realizing load balance according to an embodiment of the present invention;

[0028] Figure 2 is the structural schematic diagram of a load balancing server according to an embodiment of the present invention;

[0029] Figure 3 is the structural schematic diagram of a cluster system according to an embodiment of the present invention. Detailed Embodiments

[0030] In order to make the technical means, creative features, achieved purposes and effects realized by the present invention easy to understand, the present invention will be further described below in conjunction with specific embodiments.

[0031] Embodiment 1

[0032] Please refer to Figure 1 As shown, a method for implementing load balancing according to an embodiment of the present invention includes the following steps:

[0033] S1: According to the parameter data of the computing servers and in combination with the weighted round-robin strategy, add request allocation weights to the computing servers;

[0034] The parameter data of the computing servers in S1 includes but is not limited to memory capacity, network bandwidth, and disk read / write rate parameters;

[0035] It should be noted that the network bandwidth and disk read / write rate parameters of the computing servers are obtained through the operation reports of the computing servers in multiple historical calculation cycles, and are both the parameter averages of the computing servers in multiple historical calculation cycles; the memory capacity represents the maximum operating memory capacity of the computing servers;

[0036] The process of adding request allocation weights to the computing servers in combination with the weighted round-robin strategy includes:

[0037] Integrate the parameter data of the same type of all computing servers into parameter groups to obtain multiple parameter groups;

[0038] Based on any one computing server;

[0039] Obtain the parameter ratio of the parameter data corresponding to the computing server within the parameter group, that is, the ratio of the parameter data to the sum of the parameters within the parameter group;

[0040] Sum up all the parameter ratios corresponding to the computing server to obtain the parameter comprehensive ratio index of the computing server;

[0041] Sum up the parameter comprehensive ratio indexes of all computing servers to obtain the total parameter comprehensive ratio index;

[0042] Based on any one computing server, calculate the ratio of the parameter comprehensive ratio index of the computing server to the total parameter comprehensive ratio index to obtain the request allocation weight of the computing server;

[0043] The present application makes the following exemplary description of the process of adding request allocation weights to the computing servers:

[0044] Integrate the parameter data of the same type of all computing servers into parameter groups. Suppose there are 3 parameter groups, including:

[0045] Memory capacity groups (NC1, NC2, NC3......NCn), where NCn represents the memory capacity of the nth computing server, and n represents the serial number of the computing server;

[0046] Network bandwidth groups (WL1, WL2, WL3......WLn), where WLn represents the network bandwidth of the nth computing server, and n represents the serial number of the computing server;

[0047] Disk read / write rate parameter groups (DX1, DX2, DX3......DXn), where DXn represents the disk read / write rate of the nth computing server, and n represents the serial number of the computing server;

[0048] Based on the ith computing server, calculate the parameter ratios of the parameter data within the parameter group, including: memory capacity ratio NCB, network bandwidth ratio WLB, and disk read / write rate ratio DXB; the specific calculation is as follows:

[0049]

[0050] Sum up the memory capacity ratio NCB, network bandwidth ratio WLB, and disk read / write rate ratio DXB to obtain the parameter comprehensive ratio index ZHi of the ith computing server;

[0051] Calculate the ratio of the parameter comprehensive ratio index ZHi of the ith computing server to the total parameter comprehensive ratio index to obtain the request allocation weight of the ith computing server;

[0052] S2: Determine the ideal request amounts of each computing server in multiple historical computing cycles according to the request allocation weights, and compare and analyze them with the actual request amounts to judge whether the request loads among the computing servers are balanced;

[0053] The process of determining the ideal request amounts of each computing server in multiple historical computing cycles according to the request allocation weights is as follows:

[0054] Obtain the user request volume in each historical computing cycle, and perform a multiplication process with the request allocation weight of the computing server to obtain the ideal request amount of the computing server in the historical computing cycle;

[0055] It should be noted that the user request volume in the historical computing cycle is obtained by querying the operation logs of each computing server in the historical computing cycle;

[0056] It should be noted that the ideal request amount represents the number of user requests of the computing server calculated according to the request allocation weight of the computing server, and the actual request amount represents the actual number of user requests of the computing server in the historical computing cycle;

[0057] The process of determining whether the request loads between computing servers are balanced includes:

[0058] Compare the ideal request amount of a computing server during a historical computing cycle with the actual request amount of the computing server during the historical computing cycle;

[0059] If the ideal request amount is different from the actual request amount, mark the computing server as a non - balanced computing server;

[0060] If the ideal request amount is the same as the actual request amount, do nothing;

[0061] Based on any one historical computing cycle;

[0062] Count the proportion of the number of non - balanced computing servers in the historical computing cycle, that is, the proportion SB of the number of non - balanced computing servers in all computing servers;

[0063] Calculate the difference between the ideal request amount and the actual request amount of the non - balanced computing servers to obtain the request difference of the non - balanced computing servers. Sum up and average the request differences of all non - balanced computing servers to obtain the average request difference of non - balanced computing servers in the historical computing cycle, and perform a ratio process with the user request volume in the historical computing cycle to obtain the request difference performance value CB;

[0064] Use a geometric weighted model to calculate the load balancing value between computing servers in the historical computing cycle for the proportion SB of the number of non - balanced computing servers in all computing servers and the request difference performance value CB. Specifically:

[0065] JH = SB δ +CB γ

[0066] Where, JH is the load balancing value between computing servers in the historical computing cycle, δ and γ are weight factors, and δ + γ = 1; the weight factors are calculated by the entropy weight method;

[0067] Exemplarily, the process of calculating the weight factors by the entropy weight method is as follows:

[0068] Assume that there are b pieces of equilibrium data in the historical computing cycle, including the proportion SBz of the number of non - balanced computing servers in all computing servers and the request difference performance value CBz; z = 1, 2, 3,......, b;

[0069] Perform data standardization to standardize the data to the interval [0, 1]. The specific standardization formula is:

[0070]

[0071] Among them, x zj represents the initial data, which is the j-th index data in the z-th cycle. When j = 1, it is SB; when j = 2, it is CB; x * zj is the normalized data. min(x j ) represents the minimum value of the j-th index data, and max(x j ) represents the maximum value of the j-th index data;

[0072] Calculate the proportion:

[0073]

[0074] Calculate the entropy value:

[0075]

[0076] Calculate the coefficient of variation:

[0077] d j = 1 - e j

[0078] Calculate the weight:

[0079]

[0080] In some embodiments, the load balancing value JH between computing servers in the historical calculation period is compared with the load balancing threshold;

[0081] If the load balancing value JH is greater than or equal to the load balancing threshold, the historical calculation period is marked as an unbalanced calculation period;

[0082] If the load balancing value JH is less than the load balancing threshold, no processing is performed;

[0083] Count the proportion of the number of unbalanced calculation periods in the historical calculation period to obtain the load comprehensive value;

[0084] If the load comprehensive value is greater than or equal to the load comprehensive threshold, it indicates that there are more cases of load imbalance in the historical calculation period, which means that the load between computing servers is unbalanced;

[0085] If the load comprehensive value is less than the load comprehensive threshold, it indicates that there are fewer cases of load imbalance in the historical calculation period, which means that the load between computing servers is balanced;

[0086] S3: If it is unbalanced, predict the user request volume in the current calculation period, and combine the request allocation weights of the computing servers to determine the request allocation amount of each computing server in the current calculation period;

[0087] The process of predicting the user request volume in the current calculation period in S3 includes:

[0088] Obtain the user request volume in each historical calculation period, integrate it into a user request volume data group, process the user request volume data group using the moving average method, and predict the user request volume in the current calculation period. The specific process is as follows:

[0089] A1. Obtain the user request volume data group;

[0090] A2. Set the moving window k;

[0091] A3. Slide the moving window k, calculate the user request volume included in the moving window k, and calculate the average value. The obtained average value is used as the predicted value of the calculation period, thereby predicting the user request volume in the current calculation period;

[0092] The process of determining the request allocation of each computing server in the current calculation period in combination with the request allocation weight of the computing server is as follows:

[0093] Multiply the request allocation weight of each computing server by the user request volume in the current calculation period respectively to obtain the request allocation amount of each computing server in the current calculation period;

[0094] The technical solution of the embodiment of the present invention is: According to the parameter data of the computing server, determine the request allocation weight of the computing server, and determine the ideal request amount of each computing server in multiple historical calculation periods according to the request allocation weight, and compare and analyze it with the actual request amount to judge whether the request load between the computing servers is balanced. If it is not balanced, predict the user request volume in the current calculation period, and combine the request allocation weight of the computing server to determine the request allocation amount of each computing server in the current calculation period. The present invention realizes the judgment of whether the request load between the computing servers is balanced, and in the case of imbalance, combines the request allocation weight of the computing server and the predicted user request volume of the current calculation period to determine the request allocation amount of the computing server, ensuring the load balance of the computing server in the current calculation period.

[0095] Embodiment 2

[0096] Please refer to Figure 1 As shown, a method for realizing load balance according to the embodiment of the present invention includes the following steps:

[0097] S4: Analyze and judge the stability of the user request sending according to the request volume of the user sending requests to the computing server in multiple historical calculation periods, and determine the predicted request volume of the user in the current calculation period;

[0098] The process of analyzing and judging the stability of user request sending in S4 according to the request volume of the user sending requests to the computing server in multiple historical calculation cycles is as follows:

[0099] Obtain the request volumes of the users sending requests to the computing server in multiple historical calculation cycles, and integrate them into a user request data group (Q1, Q2, Q3......Qg), where Qg represents the request volume in the gth request sending cycle;

[0100] Calculate the standard deviation BC of the user request data group through the standard deviation formula:

[0101]

[0102] where Qi represents the request volume in the ith request sending cycle, and Q represents the mean value of the user request data group;

[0103] Perform a ratio process on the standard deviation Qi and the mean value of the user request data group to obtain the coefficient of variation BY of the user request volume. The specific calculation formula is:

[0104]

[0105] Compare the coefficient of variation BY with the coefficient of variation threshold;

[0106] If the coefficient of variation BY is greater than or equal to the coefficient of variation threshold, mark the user as an unstable request user;

[0107] If the coefficient of variation BY is less than the coefficient of variation threshold, mark the user as a stable request user;

[0108] The process of determining the predicted request volume of the user in the current calculation cycle in S4 is as follows:

[0109] If the user is a stable request user, sum up and average the request volumes of the user in multiple historical calculation cycles to obtain the predicted request volume of the user in the current calculation cycle;

[0110] If the user is an unstable request user, obtain the maximum request volume of the user in multiple historical calculation cycles as the predicted request volume of the user in the current calculation cycle. It can be understood that since the user is an unstable request user, to ensure the load balancing of the subsequent computing server, the maximum situation of the user request is considered. Therefore, the maximum request volume of the user in multiple historical calculation cycles is used as the predicted request volume of the user in the current calculation cycle;

[0111] S5: According to the predicted request volume of the user and the remaining request allocation amounts of each computing server, push the user requests to the matching computing servers in the current calculation cycle;

[0112] The process of pushing user requests to the matching computing server within the current computing cycle according to the predicted request volume of the user and the remaining request allocation amount of each computing server is as follows:

[0113] Based on any computing server, obtain the sum of the predicted request volumes of the users currently received by the computing server, and perform a difference operation with the request allocation amount of the computing server to obtain the remaining request allocation amount of the computing server;

[0114] When receiving a user's request, obtain the predicted request volume of the user;

[0115] Compare the predicted request volume of the user with the remaining request allocation amounts of each computing server respectively;

[0116] If the predicted request volume of the user is less than or equal to the remaining request allocation amount of the computing server, it means that the computing server can satisfy the request processing of the current user, and mark the computing server as a matching computing server;

[0117] If the predicted request volume of the user is greater than the remaining request allocation amount of the computing server, it means that the computing server cannot satisfy the request processing of the current user, and mark the computing server as a non-matching computing server;

[0118] Perform a difference operation between the remaining request allocation amount of the matching computing server and the predicted request volume of the user to obtain the secondary remaining request amount of the computing server; and sort the matching computing servers according to the secondary remaining request amount, and push the user request to the first matching computing server in the sorting for the processing calculation of the user request;

[0119] The present invention makes an exemplary description of the process of pushing user requests to the matching computing server;

[0120] When receiving the request of user 3, obtain the predicted request volume of user 3. Assume that the predicted request volume of user 3 is 10. Based on any computing server, for example, the users currently received requests by the computing server include user 1 and user 2; the request allocation amount of the computing server is 40, and the predicted request volumes of user 1 and user 2 are 15 and 10 respectively. Then the remaining request allocation amount of the computing server is 40 - 15 - 10 = 15;

[0121] Compare the predicted request volume 10 of user 3 with the remaining request allocation amount 15 of the computing server, that is, 10 is less than 15, which means that the computing server can satisfy the request processing of the user, and mark the computing server as a matching computing server;

[0122] Assume that there are 3 matching computing servers, namely computing server 1, computing server 2, and computing server 3;

[0123] The remaining request amounts of computing server 1, computing server 2, and computing server 3 are respectively processed by taking the difference with the predicted request volume 10 of user 3, and the secondary remaining request amounts of computing server 1, computing server 2, and computing server 3 are respectively obtained, assumed to be 5, 10, and 15; then the user request is pushed to computing server 3 for computing processing;

[0124] It should be noted that, assuming that the secondary remaining request amounts of two computing servers are the same, the user request is randomly pushed to any one of them;

[0125] The technical solution of the embodiment of the present invention is as follows: According to the request volumes of the users sending requests to the computing servers in multiple historical computing cycles, analyze and judge the stability of the user request sending, and determine the predicted request volume of the user in the current computing cycle. According to the predicted request volume of the user and the remaining request allocation amounts of each computing server, in the current computing cycle, the user is pushed to the matching computing server. The present invention realizes the matching push of user requests by analyzing the stability of user requests, predicting the request volume of users, and combining the remaining request amounts of computing servers, further realizes the request load balancing of computing servers, and pushes the user to the matching computing server according to the predicted request volume of the user, which can maintain the continuity of the user request session, improve the user experience and the logical coherence of business requests.

[0126] Embodiment 3

[0127] Please refer to Figure 2 As shown, a load balancing server described in the embodiment of the present invention includes:

[0128] Server weight calculation module: According to the parameter data of the computing server and in combination with the weighted round-robin strategy, add request allocation weights to the computing servers;

[0129] Load balancing judgment module: Determine the ideal request amounts of each computing server in multiple historical computing cycles according to the request allocation weights, and compare and analyze them with the actual request amounts to judge whether the request loads among the computing servers are balanced;

[0130] User request allocation module: If not balanced, predict the user request volume in the current computing cycle, and in combination with the request allocation weights of the computing servers, determine the request allocation amounts of each computing server in the current computing cycle;

[0131] User request prediction module: According to the request volumes of the users sending requests to the computing servers in multiple historical computing cycles, analyze and judge the stability of the user request sending, and determine the predicted request volume of the user in the current computing cycle;

[0132] User request push module: According to the predicted request volume of users and the remaining request allocation quotas of each computing server, within the current computing cycle, push user requests to the matching computing servers.

[0133] Embodiment 4

[0134] Please refer to Figure 3 As shown, a cluster system described in an embodiment of the present invention includes: a load balancing server, computing servers, and users;

[0135] The computing servers are used to receive user requests;

[0136] The load balancing server is used to add request allocation weights to the computing servers, determine whether the request loads among the computing servers are balanced, and in the case of imbalance, determine the request allocation quotas of each computing server within the current computing cycle and predict the predicted request volume of users within the current computing cycle, so as to implement the push of user requests.

[0137] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification only illustrates the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for implementing load balancing, characterized in that: include: S1: Based on the parameter data of the computing server and combined with the weighted polling strategy, add request allocation weights to the computing server; S2: Determine the ideal request amount of each computing server in multiple historical computing cycles based on the request allocation weight, and compare and analyze it with the actual request amount to determine whether the request load between computing servers is balanced; S3: If there is an imbalance, the user request volume in the current computing cycle is predicted, and the request allocation amount of each computing server in the current computing cycle is determined in combination with the request allocation weight of the computing server; S4: Analyze and determine the stability of user request sending according to the request volume of the user sending the request to the computing server in multiple historical computing cycles, and determine the predicted request volume of the user in the current computing cycle; S5: According to the predicted request volume of the user and the remaining request allocation of each computing server, the user request is pushed to the matching computing server within the current computing cycle.

2. A load balancing method according to claim 1, characterized in that: Integrate the parameter data of the same type of all computing servers into a parameter group to obtain multiple parameter groups; Obtain the parameter ratio of the parameter data corresponding to the computing server in the parameter group and sum them up to obtain the parameter comprehensive ratio index of the computing server; The parameter comprehensive proportion index of the computing server is calculated by ratio calculation with the parameter comprehensive proportion total index to obtain the computing server request allocation weight, wherein the parameter comprehensive proportion total index is the sum of the parameter comprehensive proportion indices of all computing servers.

3. A load balancing method according to claim 1, characterized in that: The user request amount in each historical computing cycle is obtained and multiplied by the request allocation weight of the computing server to obtain the ideal request amount of the computing server in the historical computing cycle.

4. A load balancing method according to claim 1, characterized in that: If the ideal request amount is different from the actual request amount, the computing server is marked as an unbalanced computing server; By analyzing the number of unbalanced computing servers and the deviation between the corresponding ideal request amount and the actual request amount, the proportion of unbalanced computing servers in all computing servers and the request difference performance value are obtained, and the load balancing value between computing servers in the historical computing cycle is obtained by processing using a geometric weighted model; If the load balancing value is greater than or equal to the load balancing threshold, the historical calculation period is marked as an unbalanced calculation period; Count the proportion of unbalanced computing cycles in historical computing cycles to obtain the comprehensive load value; If the combined load value is greater than or equal to the combined load threshold, it indicates that the loads between computing servers are unbalanced.

5. A load balancing method according to claim 4, characterized in that: The difference between the ideal request amount and the actual request amount of the unbalanced computing server is calculated to obtain the request difference of the unbalanced computing server. The request differences of all unbalanced computing servers are summed and averaged, and then compared with the user request volume in the historical computing cycle to obtain the request difference performance value.

6. A load balancing method according to claim 1, characterized in that: Obtain the user request volume in each historical calculation cycle and integrate it into a user request volume data group. Use the moving average method to process the user request volume data group and predict the user request volume in the current calculation cycle. The request allocation weight of each computing server is multiplied by the user request volume in the current computing cycle to obtain the request allocation amount of each computing server in the current computing cycle.

7. A load balancing method according to claim 1, characterized in that: Obtain the request volume of users who send requests to the computing server in multiple historical computing cycles and integrate them into a user request data group. Calculate the coefficient of variation based on the user request data group, mark users whose coefficient of variation is greater than or equal to a coefficient of variation threshold as unstable request users, and mark users whose coefficient of variation is less than the coefficient of variation threshold as stable request users. If the user is a stable request user, the user's request volume in multiple historical computing cycles is summed and averaged to obtain the predicted request volume of the user in the current computing cycle; If the user is a non-stable request user, the maximum request amount of the user in multiple historical computing cycles is obtained as the predicted request amount of the user in the current computing cycle.

8. A load balancing method according to claim 1, characterized in that: Obtain the sum of the predicted requests currently received by the computing server, and perform subtraction processing on the request allocation of the computing server to obtain the remaining request allocation of the computing server; If the user's predicted request volume is less than or equal to the remaining request allocation of the computing server, the computing server is marked as a matching computing server; If the user's predicted request volume is greater than the remaining request allocation of the computing server, the computing server will be marked as a non-matching computing server; The remaining request allocation of the matching computing server is processed with the difference between the predicted request amount of the user to obtain the secondary remaining request amount of the computing server; The matching calculation servers are sorted according to the secondary residual request amount, and the user request is pushed to the matching calculation server ranked first to process and calculate the user request.

9. A load balancing server, characterized in that: include: Server weight calculation module: Based on the parameter data of the computing server and combined with the weighted polling strategy, it adds request allocation weights to the computing server; Load balancing judgment module: Determines the ideal request amount of each computing server in multiple historical computing cycles based on the request allocation weight, and compares and analyzes it with the actual request amount to determine whether the request load between computing servers is balanced; User request allocation module: If there is an imbalance, the user request volume in the current computing cycle is predicted, and the request allocation amount of each computing server in the current computing cycle is determined in combination with the request allocation weight of the computing server; User request prediction module: Analyzes and determines the stability of user request sending based on the request volume of users sending requests to the computing server in multiple historical computing cycles, and determines the predicted request volume of users in the current computing cycle; User request push module: pushes user requests to matching computing servers within the current computing cycle based on the user's predicted request volume and the remaining request allocation of each computing server.

10. A cluster system, characterized in that: include: Load balancing servers, computing servers, and users; The computing server is used to receive user requests; The load balancing server is used to add request allocation weights to the computing servers, determine whether the request loads between the computing servers are balanced, and in the case of imbalance, determine the request allocation quota of each computing server in the current computing cycle and predict the predicted request volume of the user in the current computing cycle to realize user request push.