Server dynamic load balancing and resource management method
By monitoring server CPU utilization in real time and analyzing historical data, server grouping and task allocation are dynamically adjusted, solving the problem of unreasonable server resource allocation and realizing intelligent management and efficient utilization of server resources.
Patent Information
- Application Number
- CN202511312893.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-15
- Publication Date
- 2025-12-12
AI Technical Summary
Existing server load balancing and resource management methods lack the ability to dynamically perceive and adaptively adjust the real-time operating status of servers, resulting in unreasonable allocation of server resources, with some servers overloaded while others are idle, affecting overall performance and resource utilization.
By monitoring server CPU utilization in real time, dynamically adjusting server grouping and task allocation strategies, and combining historical data to analyze future load fluctuation trends, intelligent management of server resources is achieved. A distributed probe cluster is used to collect CPU utilization in real time at the second level, and clear CPU utilization thresholds and load fluctuation judgment mechanisms are set to perform elastic scaling operations.
It enables precise perception of server operating status and timely resource adjustment, effectively avoiding server overload or idleness and improving the overall utilization rate of resources.
Smart Images

Figure CN121116636A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of server load balancing technology, and in particular to a method for dynamic server load balancing and resource management. Background Technology
[0002] With the rapid development of information technology, servers are increasingly widely used in various fields, becoming the core infrastructure supporting the stable operation of various business systems. The efficient operation of servers is of great significance for ensuring business continuity, improving user experience, and reducing operating costs. Therefore, in order to ensure the efficient operation of servers, server load balancing and resource management methods are used to manage servers.
[0003] However, existing server load balancing and resource management methods have significant drawbacks. Most of them lack the ability to dynamically perceive and adaptively adjust the real-time operating status of servers. They cannot timely and reasonably group and allocate tasks based on the actual CPU utilization of servers, nor can they elastically scale up or down servers according to load fluctuations in future periods. This can easily lead to unreasonable allocation of server resources, with some servers overloaded while others are idle, affecting the overall performance and resource utilization of the servers. Summary of the Invention
[0004] (a) Purpose of the invention
[0005] To address the technical problems existing in the background art, this invention proposes a dynamic load balancing and resource management method for servers. By monitoring the server CPU utilization in real time, dynamically adjusting server grouping and task allocation strategies, and combining historical data to analyze future load fluctuation trends, intelligent management of server resources is achieved. This method can more accurately perceive the server's operating status, make timely resource adjustments, effectively avoid the problem of some servers being overloaded while others are idle, and improve the overall utilization rate of server resources.
[0006] (II) Technical Solution
[0007] This invention provides a method for dynamic load balancing and resource management of servers, specifically including the following steps:
[0008] S1: Obtain the CPU utilization of multiple servers in real time, group the multiple servers according to the obtained CPU utilization, divide the multiple servers into priority queue and idle queue, and arrange the multiple servers in the priority queue in order of CPU utilization from low to high.
[0009] S2: Retrieve real-time service requests and execute them according to the priority queue.
[0010] S3: Obtain the load volume for the next half hour of the same period within the past week. Compare the obtained load volumes. If the increase in load volume for the next half hour of the same period within the past week is greater than 30%, then there is fluctuation. If the increase in load volume for the next half hour of the same period within the past week is less than 30%, then there is no fluctuation. Perform elastic scaling up and down operations on the server based on the obtained load fluctuations.
[0011] Preferably, the specific steps for grouping multiple servers are as follows: the CPU utilization of the acquired multiple servers is judged. If the CPU utilization is greater than 30%, it is added to the priority queue; if the CPU utilization is less than 30%, it is added to the idle queue.
[0012] Preferably, the specific steps for elastic scaling up and down of servers are as follows: determine the servers in the priority queue and perform scaling up based on the determination structure; determine the servers in the idle queue and perform scaling down based on the determination structure.
[0013] Preferably, the servers in the priority queue are judged as follows: obtain the CPU utilization of the servers in the priority queue in the past half hour. If the obtained CPU utilization of the server in the past half hour is consistently greater than 50% and fluctuates, then the server is expanded. If the obtained CPU utilization of the server in the past half hour is consistently greater than 50% and does not fluctuate, then no operation is performed on the server. If the obtained CPU utilization of the server in the past half hour is less than 50%, no corresponding operation is performed on the server regardless of whether there is fluctuation.
[0014] Preferably, the specific steps for judging the servers in the idle queue are as follows: obtain the CPU utilization of the servers in the idle queue in the past half hour. If the CPU utilization of the obtained servers in the past half hour is consistently less than 30% and there is no fluctuation, then the server is scaled down. If the CPU utilization of the obtained servers in the past half hour is consistently less than 30% and there is fluctuation, then no operation is performed on the server. If the CPU utilization of the obtained servers in the past half hour is greater than 30%, no corresponding operation is performed on the server regardless of whether there is fluctuation.
[0015] Preferably, expanding the server capacity specifically includes the following steps: adding more nodes to the server.
[0016] Preferably, the server downsizing operation specifically includes the following steps: migrating data to the server in batches.
[0017] Preferably, the specific steps for obtaining the CPU utilization of multiple servers in real time include collecting the CPU utilization of the server cluster in real time with a granularity of seconds through a distributed probe cluster deployed on a Kubernetes DaemonSet.
[0018] Preferably, the specific steps for obtaining the CPU utilization of multiple servers in real time include: The specific calculation formula for real-time collection of the CPU utilization of the server cluster at a second-level granularity using a distributed probe cluster deployed on a Kubernetes DaemonSet is as follows: ,in Let be the CPU utilization at time t. This represents the CPU's idle time at time t. This represents the sum of the CPU time field at time t.
[0019] Compared with the prior art, the above-mentioned technical solution of the present invention has the following beneficial technical effects:
[0020] In this invention, the technical solution achieves intelligent management of server resources by monitoring server CPU utilization in real time, dynamically adjusting server grouping and task allocation strategies, and combining historical data analysis to predict future load fluctuation trends. This method can more accurately perceive the server's operating status, make timely resource adjustments, effectively avoid the problem of some servers being overloaded while others are idle, and improve the overall utilization rate of server resources. Attached Figure Description
[0021] Figure 1 This is a flowchart of a server dynamic load balancing and resource management method proposed in this invention. Detailed Implementation
[0022] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to specific embodiments and the accompanying drawings. It should be understood that these descriptions are merely exemplary and not intended to limit the scope of the invention. Furthermore, descriptions of well-known structures and techniques are omitted in the following description to avoid unnecessarily obscuring the concept of the invention.
[0023] In the description of the invention, it should be noted that the terms "upper," "lower," "inner," "outer," "front end," "rear end," "both ends," "one end," and "the other end," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are used only for the convenience of describing the invention and for simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the invention. Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.
[0024] In the description of the invention, it should be noted that, unless otherwise explicitly specified and limited, the terms "installed," "equipped with," and "connected," etc., should be interpreted broadly. For example, "connected" can be a fixed connection, such as welding, riveting, or bonding; it can also be a detachable connection, such as threaded connection, keyed connection, or pin connection; or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium; or it can be a connection within two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.
[0025] Example 1:
[0026] like Figure 1 As shown, the server dynamic load balancing and resource management method proposed in this invention specifically includes the following steps:
[0027] The CPU utilization of multiple servers is acquired in real time. Based on the acquired CPU utilization, the servers are grouped into a priority queue and an idle queue. The servers in the priority queue are then arranged in ascending order of CPU utilization. The CPU utilization can be collected in real time at the second level by a distributed probe cluster deployed on a Kubernetes DaemonSet.
[0028] Retrieve real-time service requests and execute them according to the priority queue. Specifically, service requests can be distributed in the queue in ascending order of server CPU utilization to ensure load balancing.
[0029] Analyze load fluctuations for future periods and perform elastic scaling up or down operations on the servers based on the obtained load fluctuations.
[0030] This technical solution achieves intelligent management of server resources by monitoring server CPU utilization in real time, dynamically adjusting server grouping and task allocation strategies, and combining historical data analysis to predict future load fluctuation trends. This method can more accurately perceive the server's operating status, make timely resource adjustments, effectively avoid the problem of some servers being overloaded while others are idle, and improve the overall utilization rate of server resources.
[0031] Furthermore, this application also proposes specific steps for grouping multiple servers: judging the CPU utilization of the acquired multiple servers, if the CPU utilization is greater than 30%, then adding them to the priority queue, and if the CPU utilization is less than 30%, then adding them to the idle queue.
[0032] Therefore, this technical solution achieves automated classification and management of server resources by setting a clear CPU utilization threshold standard. The 30% threshold setting is based on a large amount of experimental data verification and can effectively distinguish between high-load and low-load server states. Through quantitative indicators and automated judgment, the timeliness and accuracy of server grouping are ensured, laying a reliable foundation for subsequent load balancing and resource scheduling.
[0033] Furthermore, this application also proposes specific steps for analyzing load fluctuations in future periods. By obtaining the load volume in the next half hour of the same period in the past week, it is determined whether there are fluctuations based on the obtained load volume. The load volume can be obtained through historical monitoring data or log analysis. For example, a time series database can be used to store historical server load data, and the load volume for a specified period can be obtained through a query interface.
[0034] Therefore, this technical solution analyzes future load fluctuations using historical data, providing a basis for decision-making in subsequent elastic scaling operations. This solution can predict load change trends based on historical data, avoiding the limitations of relying solely on the current load status for resource adjustments. By analyzing load data from the same period last week, it can identify periodic or sudden load change patterns, thereby more accurately determining whether resource adjustments need to be made in advance.
[0035] Furthermore, this application also proposes specific steps for determining the existence of load fluctuations. By comparing the load volume of the next half hour in the same period of the past week, it is determined that there is a fluctuation when the load volume increases by more than 30%, and that there is no fluctuation when the increase is less than 30%. That is, the increase is calculated by using the maximum and minimum load volume of the next half hour in the same period of the past week.
[0036] This technical solution determines load fluctuation status by quantifying thresholds, solving the problem that traditional methods cannot objectively identify sudden load changes. The fluctuation determination mechanism established thereby provides a reliable basis for subsequent flexible scaling decisions. The 30% threshold setting considers both system fault tolerance requirements and timely captures significant load changes. In practice, historical load data can be stored in a time-series database, and the percentage increase can be calculated in real time using a streaming computing framework.
[0037] Furthermore, this application also proposes the following specific steps for performing elastic scaling up and down operations on servers: servers in the priority queue are assessed, and scaling up is performed based on the assessment structure; servers in the idle queue are assessed, and scaling down is performed based on the assessment structure.
[0038] Therefore, this technical solution achieves precise scaling up and down control of server resources by establishing a dual judgment mechanism based on real-time CPU utilization and historical load fluctuations. This solution can dynamically adjust resource allocation according to the actual operating status of the server, avoiding resource waste and performance bottlenecks. By dividing priority queues and idle queues and combining historical load fluctuation analysis, it can accurately identify high-load servers that need to be expanded and low-load servers that are suitable for scaling down, thereby optimizing the overall resource utilization.
[0039] Furthermore, this application proposes a specific method for judging servers in the priority queue. First, it obtains the CPU utilization data of servers in the priority queue over the past half hour. If the CPU utilization of a server has consistently exceeded 50% over the past half hour and there are load fluctuations, then a capacity expansion operation is performed on that server. If the CPU utilization of a server has consistently exceeded 50% over the past half hour but there are no load fluctuations, then the current state is maintained without adjustment. In addition, if the CPU utilization of a server has fallen below 50% over the past half hour, regardless of whether there are load fluctuations, no capacity expansion operation is performed on that server.
[0040] Furthermore, this application also proposes specific steps for judging servers in the idle queue: obtain the CPU utilization of servers in the idle queue over the past half hour; if the CPU utilization of the server over the past half hour is consistently less than 30% and there is no fluctuation, then perform a scaling-down operation on the server; if the CPU utilization is consistently less than 30% but there is fluctuation, then no operation is performed on the server; if the CPU utilization is greater than 30%, no operation is performed regardless of whether there is fluctuation.
[0041] Therefore, this technical solution effectively avoids the risk of accidental scaling down by using a dual-condition judgment mechanism: first, low-load servers are screened based on real-time utilization data, and then historical fluctuation analysis is combined to ensure load stability.
[0042] Furthermore, this application also proposes that expanding the server capacity specifically includes the following steps: adding more nodes to the server.
[0043] Specifically, adding more nodes refers to improving the processing capacity of the server cluster through horizontal scaling. This technical solution achieves precise matching between resource supply and load demand through a modular node expansion mechanism. When servers in the priority queue are continuously under high load and fluctuate, the expansion process is automatically triggered, avoiding performance degradation due to insufficient resources.
[0044] Furthermore, this application also proposes specific steps for scaling down the server, including migrating data to the server in batches.
[0045] This technical solution achieves smooth server scaling down through batch migration, avoiding performance fluctuations or service interruptions caused by migrating a large amount of data at once. This method can gradually release idle server resources and improve resource utilization without affecting business continuity.
[0046] Furthermore, this application also proposes specific steps for obtaining the CPU utilization of multiple servers in real time, by using a distributed probe cluster deployed on a Kubernetes DaemonSet to collect the CPU utilization of the server cluster in real time at a granular level of seconds.
[0047] Furthermore, the specific steps for obtaining the CPU utilization of multiple servers in real time include the following: The specific calculation formula for real-time collection of the CPU utilization of the server cluster at a second-level granularity using a distributed probe cluster deployed on a Kubernetes DaemonSet is as follows:
[0048] ,in Let be the CPU utilization at time t. This represents the CPU's idle time at time t. This is the sum of CPU time fields at time t, including user, nice, system, idle, iowait, irq, and softirq. Where: user: CPU time executing processes in user mode (unit: jiffies, typically 1 jiffies = 10ms); nice: CPU time executing low-priority (nice value > 0) processes in user mode; system: CPU time executing system calls in kernel mode; idle: CPU idle time (excluding I / O wait); iowait: CPU idle time with outstanding I / O requests; irq: CPU time handling hardware interrupts; softirq: CPU time handling software interrupts; steal: time stolen by other virtual machines in a virtualized environment; guest: time running virtual guests; guest_nice: time running low-priority virtual guests.
[0049] Two were implemented;
[0050] When the distributed probe cluster deployed on Kubernetes DaemonSet collects real-time CPU utilization data of the server cluster at a granular level every 5 seconds, the first sampled data is: cpu 1000, 50, 300, 5000, 100, 10, 200, 0, 0.
[0051] but =1000+50+300+5000+100+10+20=6480,
[0052] The second sampled data was: CPU 5200, 250, 1500, 24000, 500, 50, 100, 0, 0 / 0
[0053] but =5200+250+1500+24000+500+50+100=31600,
[0054] =25120, =24000−5000=19000,
[0055]
[0056] Therefore, the CPU utilization rate is 24.4%.
[0057] It should be understood that the specific embodiments described above are merely illustrative or explanatory of the principles of the invention and do not constitute a limitation thereof. Therefore, any modifications, equivalent substitutions, improvements, etc., made without departing from the spirit and scope of the invention should be included within the protection scope of the invention. Furthermore, the appended claims are intended to cover all variations and modifications falling within the scope and boundaries of the appended claims, or equivalent forms of such scope and boundaries.
Claims
1. A method for dynamic load balancing and resource management of servers, characterized in that, Specifically, the following steps are included: S1: Real-time acquisition of CPU utilization of multiple servers, grouping multiple servers according to the acquired CPU utilization of multiple servers, dividing multiple servers into priority queue and idle queue, and arranging multiple servers in the priority queue in order of CPU utilization from low to high. S2: Obtain real-time service requests and execute them according to the order of the priority queue. S3: Obtain the load volume in the next half hour of the same period within the past week, compare the obtained load volume, if the increase in the load volume in the next half hour of the same period within the past week is greater than 30%, then there is fluctuation, if the increase in the load volume in the next half hour of the same period within the past week is less than 30%, then there is no fluctuation, and perform elastic scaling up and down operation on the server according to the obtained load fluctuation.
2. The server dynamic load balancing and resource management method according to claim 1, characterized in that, The specific steps for grouping multiple servers are as follows: the CPU utilization of the acquired servers is judged. If the CPU utilization is greater than 30%, the server is added to the priority queue. If the CPU utilization is less than 30%, the server is added to the idle queue.
3. The server dynamic load balancing and resource management method according to claim 1, characterized in that, The specific steps for elastic scaling up and down of the server are as follows: the server in the priority queue is judged, and the scaling up operation is performed according to the judgment structure; the server in the idle queue is judged, and the scaling down operation is performed according to the judgment structure.
4. The server dynamic load balancing and resource management method according to claim 3, characterized in that, The servers in the priority queue are judged as follows: the CPU utilization of the servers in the priority queue over the past half hour is obtained. If the CPU utilization of the servers over the past half hour is consistently greater than 50% and fluctuates, then the server is scaled up. If the CPU utilization of the servers over the past half hour is consistently greater than 50% and does not fluctuate, then no operation is performed on the server. If the CPU utilization of the servers over the past half hour is less than 50%, no corresponding operation is performed on the server regardless of whether there is fluctuation.
5. The server dynamic load balancing and resource management method according to claim 4, characterized in that, The specific steps for judging the servers in the idle queue are as follows: obtain the CPU utilization of the servers in the idle queue over the past half hour. If the CPU utilization of the servers over the past half hour is consistently less than 30% and there is no fluctuation, then the servers are scaled down. If the CPU utilization of the servers over the past half hour is consistently less than 30% and there is fluctuation, then no operation is performed on the servers. If the CPU utilization of the servers over the past half hour is greater than 30%, no corresponding operation is performed on the servers regardless of whether there is fluctuation.
6. The server dynamic load balancing and resource management method according to claim 4, characterized in that, The expansion operation of the server specifically includes the following steps: adding more nodes to the server.
7. The server dynamic load balancing and resource management method according to claim 5, characterized in that, The specific steps for scaling down the server include: migrating data to the server in batches.
8. The server dynamic load balancing and resource management method according to claim 1, characterized in that, The specific steps for obtaining the CPU utilization of multiple servers in real time include collecting the CPU utilization of the server cluster in real time with a granularity of seconds through a distributed probe cluster deployed on a Kubernetes DaemonSet.
9. A server dynamic load balancing and resource management method according to claim 8, characterized in that, The specific steps for obtaining the CPU utilization of multiple servers in real time include: using a distributed probe cluster deployed on a Kubernetes DaemonSet to collect CPU utilization of the server cluster in real time at a granular level (second-level). The specific calculation formula is as follows: ,in Let be the CPU utilization at time t. This represents the CPU's idle time at time t. This represents the sum of the CPU time field at time t.
Citation Information
Cited By
Data stream processing method and device for binding kernel and network card queue, and medium
CN121935031A