Load balancing management system for multiple display cards of server
The server multi-GPU load balancing management system, which monitors GPU parameters in real time and schedules tasks based on preset algorithms, solves the problems of GPU overload and heterogeneous GPU task matching, realizes intelligent load balancing and resource optimization of GPU clusters, and improves system stability and computing efficiency.
Patent Information
- Application Number
- CN202511448334.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-11
- Publication Date
- 2026-01-02
AI Technical Summary
Existing server multi-GPU systems lack a dynamic sensing and adjustment mechanism for the real-time load status of GPUs, leading to GPU overload. Furthermore, intelligent task matching is difficult to achieve in heterogeneous environments where different models of GPUs are used together.
The system employs a load acquisition module to monitor graphics card parameters in real time, a central scheduling module to schedule tasks based on a preset algorithm, a graphics card control module to perform task allocation and migration, and a status feedback module to form a closed-loop control, thereby achieving graphics card load balancing.
It achieves intelligent load balancing and resource optimization for heterogeneous graphics card clusters, improving system stability and computing efficiency.
Smart Images

Figure CN121255463A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of server hardware resource management technology, specifically a server multi-GPU load balancing management system. Background Technology
[0002] In fields such as high-performance computing, artificial intelligence training, deep learning inference, and graphics rendering, servers are typically configured with multiple high-performance graphics cards to meet the needs of large-scale parallel computing.
[0003] However, most existing servers use a static task allocation method, which assigns tasks to specific graphics cards in a fixed manner. This lacks a dynamic perception and adjustment mechanism for the real-time load status of the graphics cards. When the computing power or video memory requirements of a certain task suddenly increase, the corresponding graphics card is prone to overload.
[0004] Existing multi-GPU load balancing management systems generally suffer from the following problems: 1. Relying solely on a single indicator such as computing power utilization cannot fully reflect the true state of a graphics card; 2. In heterogeneous environments where different models of graphics cards are used together, traditional methods are difficult to achieve intelligent task matching.
[0005] Therefore, this application proposes a server multi-GPU load balancing management system. Summary of the Invention
[0006] The purpose of this invention is to provide a server multi-GPU load balancing management system to solve the above-mentioned technical problems: The objective of this invention can be achieved through the following technical solutions: A server multi-GPU load balancing management system, the system comprising a load acquisition module, a central scheduling module, a GPU control module, and a status feedback module; The load acquisition module is used to establish data communication connections with each graphics card in the server and collect the load parameters of each graphics card in real time. The load parameters include core temperature, memory usage, computing power utilization and task queue length. The central scheduling module is communicatively connected to the load acquisition module, the graphics card control module, and the status feedback module, respectively. It is used to receive the load parameters transmitted by the load acquisition module, perform comprehensive analysis on the load parameters based on a preset load balancing algorithm, and generate task scheduling instructions. The graphics card control module is communicatively connected to the central scheduling module and each graphics card, and is used to receive task scheduling instructions sent by the central scheduling module, and to perform task allocation and migration according to the task scheduling instructions; The status feedback module is communicatively connected to the graphics card control module and each graphics card. It is used to collect the real-time operating status parameters of each graphics card after load adjustment and to send the real-time operating status parameters back to the central scheduling module to form a load balancing closed-loop control.
[0007] As a further description of the technical solution of the present invention, the load acquisition module includes a parameter acquisition unit and a data preprocessing unit; The working process of the parameter acquisition unit includes: numbering all the graphics cards on the server, with the numbers being 1, 2, ..., n in sequence; Collect the temperature, memory usage, computing power utilization, and task queue length of each of the n graphics card cores. The data preprocessing unit's operation includes: noise reduction, outlier removal, and standardization of the collected load parameters, wherein outlier removal employs a 3D algorithm. The standardization process uses a normalization method to convert load parameters to the same standard.
[0008] As a further description of the technical solution of the present invention, the working process of the central scheduling module includes: Obtain the average values of various load parameters for the i-th graphics card during the peak hours of the server's current work cycle. Based on the obtained average values of the parameters, construct a mathematical model for the first state coefficient of the i-th graphics card, expressed as:
[0009] In the formula, Let be the average core temperature during the peak period of the current working cycle of the i-th graphics card. The upper limit temperature set for the i-th peak graphics card time period of the system. Let be the average video memory usage during peak hours in the current working cycle of the i-th graphics card. Let be the computing power utilization rate of the i-th graphics card during the peak period of the current working cycle. Let be the length of the task queue during peak hours in the current working cycle of the i-th graphics card. This is the first state coefficient of the graphics card.
[0010] As a further description of the technical solution of the present invention, the working process of the central scheduling module also includes: Obtain the average values of various load parameters for the i-th graphics card during the server's off-peak hours in the current work cycle. Based on the obtained average parameter values, construct a mathematical model for the second state coefficient of the i-th graphics card, expressed as:
[0011] In the formula, Let be the average core temperature during the lowest point of the current operating cycle of the i-th graphics card. The upper limit temperature set for the i-th low-frequency period of the graphics card in the system. Let be the average video memory usage during the off-peak period in the current working cycle of the i-th graphics card. Let be the computing power utilization rate of the i-th graphics card during the off-peak period in the current working cycle. Let be the length of the task queue during the off-peak period in the current working cycle of the i-th graphics card. This is the second state coefficient of the graphics card.
[0012] As a further description of the technical solution of the present invention, the working process of the central scheduling module also includes: Based on the first and second state coefficients of the i-th graphics card, a mathematical model of the comprehensive state coefficients of the i-th graphics card is constructed, and its expression is: ; In the formula, Let i be the overall state coefficient of the i-th graphics card. Compared with the first preset threshold set by the system If a comparison is made, Less than If the i-th graphics card is in an unsatisfactory working state, then the tasks of the i-th graphics card need to be transferred to other graphics cards.
[0013] As a further description of the technical solution of the present invention, the method for migrating the working task of the i-th graphics card to other graphics cards includes: Calculate the overall status coefficients of n graphics cards sequentially, and then compare the overall status coefficients of the n tabs with the second preset threshold set by the system. In comparison, among them, < Filter out graphics cards whose overall performance coefficient is greater than or equal to the second preset threshold. All m graphics cards are migration objects, where m belongs to n; From the m graphics cards, further select compatible cards for migration.
[0014] As a further description of the technical solution of the present invention, the method for further selecting an adapter card for migration from m graphics cards includes: Obtain the peak computing power, memory bandwidth, and memory capacity of the j-th graphics card, where j belongs to m. Construct a mathematical model for the ji migration index coefficient, with the expression: ; In the formula, , These are the peak computing power of the j-th and i-th graphics cards, respectively. , These are the video memory bandwidths of the j-th and i-th graphics cards, respectively. , Let be the video memory capacities of the j-th and i-th graphics cards, respectively. Let be the migration index coefficient of the j-th graphics card relative to the i-th graphics card; like If the value is approximately equal to 1, it means the two cards have similar performance and can be used interchangeably; if... If the value is greater than 1, it means that j is significantly stronger than i; if If the value is much greater than 1, it means that j is significantly stronger than i, and the performance may be excessive. Calculate the migration index coefficients of m graphics cards relative to the i-th graphics card in sequence, and select the graphics card with a migration index coefficient greater than 1 and closest to 1 as the i-th graphics card task migration object.
[0015] As a further description of the technical solution of the present invention, the status feedback module transmits the adjusted graphics card running status back to the central scheduling module in real time, dynamically adjusts the scheduling strategy, and ensures the stability of the balancing effect.
[0016] The beneficial effects of this invention are: The system of this invention includes a load acquisition module, a central scheduling module, a graphics card control module, and a status feedback module. The load acquisition module collects real-time data on the core temperature, memory usage, computing power utilization, and task queue length of each graphics card. Based on the collected parameters, the central scheduling module constructs a status coefficient model for peak and off-peak periods, calculates a comprehensive status coefficient to assess the health of the graphics cards, and uses migration index coefficients to accurately select target graphics cards. The graphics card control module performs task allocation and migration. The status feedback module transmits the adjusted status parameters back in real-time, forming a closed-loop control. This invention achieves intelligent load balancing and resource optimization for heterogeneous graphics card clusters through nonlinear modeling and multi-dimensional evaluation, effectively improving system stability and computing efficiency. Attached Figure Description
[0017] The invention will now be further described with reference to the accompanying drawings.
[0018] Figure 1 This is a schematic diagram of the server multi-GPU load balancing management system of the present invention. Detailed Implementation
[0019] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0020] Please see Figure 1As shown, a server multi-GPU load balancing management system includes a load acquisition module, a central scheduling module, a GPU control module, and a status feedback module. The load acquisition module is used to establish data communication connections with each graphics card in the server and collect the load parameters of each graphics card in real time. The load parameters include core temperature, memory usage, computing power utilization and task queue length. The central scheduling module is communicatively connected to the load acquisition module, the graphics card control module, and the status feedback module, respectively. It is used to receive the load parameters transmitted by the load acquisition module, perform comprehensive analysis on the load parameters based on a preset load balancing algorithm, and generate task scheduling instructions. The graphics card control module is communicatively connected to the central scheduling module and each graphics card, and is used to receive task scheduling instructions sent by the central scheduling module, and to perform task allocation and migration according to the task scheduling instructions; The status feedback module is communicatively connected to the graphics card control module and each graphics card. It is used to collect the real-time operating status parameters of each graphics card after load adjustment and to send the real-time operating status parameters back to the central scheduling module to form a load balancing closed-loop control.
[0021] Through the above technical solution, this invention provides a closed-loop control process that dynamically monitors, analyzes, and adjusts the workload of multiple graphics cards in a server to achieve optimal overall performance and hardware health. The load acquisition module assigns a unique identifier to each graphics card and continuously collects four key load parameters: core temperature, memory usage, computing power utilization, and task queue length. It also cleans and standardizes the collected raw data (e.g., noise reduction, outlier removal, and normalization) to provide high-quality, consistent input for subsequent analysis. The central scheduling module calculates the state coefficients of each graphics card during peak and off-peak hours, then combines these coefficients into a harmonic average to comprehensively reflect the health status of the graphics cards throughout the entire work cycle, thus determining if a particular graphics card is unhealthy. The system needs to migrate the tasks of unhealthy graphics cards. It selects the best target graphics card from among other healthy ones, not necessarily the most powerful, but one with a replacement index greater than or closest to 1, to achieve a smooth performance replacement and avoid resource waste and excessive migration costs. The graphics card control module receives task scheduling instructions from the central scheduling module and precisely executes task allocation or migration operations, transferring the computing load from the unhealthy graphics card to the selected healthy target graphics card. After load adjustment, the system collects real-time operating status parameters of each graphics card again and sends this new status data back to the central scheduling module to form a closed-loop control. The central scheduling module can verify the effectiveness of the scheduling strategy based on the feedback results and dynamically adjust subsequent decisions, enabling the system to continuously adapt to changing workloads.
[0022] As a further description of the technical solution of the present invention, the load acquisition module includes a parameter acquisition unit and a data preprocessing unit; The working process of the parameter acquisition unit includes: numbering all the graphics cards on the server, with the numbers being 1, 2, ..., n in sequence; Collect the temperature, memory usage, computing power utilization, and task queue length of each of the n graphics card cores. The data preprocessing unit's operation includes: noise reduction, outlier removal, and standardization of the collected load parameters, wherein outlier removal employs a 3D algorithm. The standardization process uses a normalization method to convert load parameters to the same standard.
[0023] As a further description of the technical solution of the present invention, the working process of the central scheduling module includes: Obtain the average values of various load parameters for the i-th graphics card during the peak hours of the server's current work cycle. Based on the obtained average values of the parameters, construct a mathematical model for the first state coefficient of the i-th graphics card, expressed as:
[0024] In the formula, Let be the average core temperature during the peak period of the current working cycle of the i-th graphics card. The upper limit temperature set for the i-th peak graphics card time period of the system. Let be the average video memory usage during peak hours in the current working cycle of the i-th graphics card. Let be the computing power utilization rate of the i-th graphics card during the peak period of the current working cycle. Let be the length of the task queue during peak hours in the current working cycle of the i-th graphics card. This is the first state coefficient of the graphics card.
[0025] Through the above technical solution, this embodiment defines how to quantify the health status or load capacity of a graphics card during peak hours. By using a multiplicative relationship, it integrates four key parameters to construct a mathematical model of the first state coefficient of the i-th graphics card. In the formula, For temperature, the higher the temperature (especially above), the better. Afterwards, this value drops sharply, thus significantly lowering the overall state coefficient. This is the memory utilization rate. As the utilization rate increases, this value gradually decreases, and it decreases very rapidly when the utilization rate is extremely high, reflecting the severity of the memory bottleneck. This is a task efficiency item, which evaluates the efficiency of the graphics card in processing tasks. This is an efficiency metric that measures the amount of tasks accumulated per unit of computing power utilization. The closer the value is to 1, the higher the efficiency of the graphics card. Conversely, it indicates that the system is inefficient. This embodiment provides an accurate and reliable mathematical basis for intelligent load migration decisions.
[0026] As a further description of the technical solution of the present invention, the working process of the central scheduling module also includes: Obtain the average values of various load parameters for the i-th graphics card during the server's off-peak hours in the current work cycle. Based on the obtained average parameter values, construct a mathematical model for the second state coefficient of the i-th graphics card, expressed as:
[0027] In the formula, Let be the average core temperature during the lowest point of the current operating cycle of the i-th graphics card. The upper limit temperature set for the i-th low-frequency period of the graphics card in the system. Let be the average video memory usage during the off-peak period in the current working cycle of the i-th graphics card. Let be the computing power utilization rate of the i-th graphics card during the off-peak period in the current working cycle. Let be the length of the task queue during the off-peak period in the current working cycle of the i-th graphics card. This is the second state coefficient of the graphics card.
[0028] Through the above technical solution, this embodiment constructs a mathematical model for evaluating the state of a graphics card during off-peak hours. The working principle is consistent with how to quantify the health status or load capacity of a graphics card during peak hours.
[0029] As a further description of the technical solution of the present invention, the working process of the central scheduling module also includes: Based on the first and second state coefficients of the i-th graphics card, a mathematical model of the comprehensive state coefficients of the i-th graphics card is constructed, and its expression is: ; In the formula, Let i be the overall state coefficient of the i-th graphics card. Compared with the first preset threshold set by the system If a comparison is made, Less than If the i-th graphics card is in an unsatisfactory working state, then the tasks of the i-th graphics card need to be transferred to other graphics cards.
[0030] As a further description of the technical solution of the present invention, the method for migrating the working task of the i-th graphics card to other graphics cards includes: Calculate the overall status coefficients of n graphics cards sequentially, and then compare the overall status coefficients of the n tabs with the second preset threshold set by the system. In comparison, among them, < Filter out graphics cards whose overall performance coefficient is greater than or equal to the second preset threshold. All m graphics cards are migration objects, where m belongs to n; From the m graphics cards, further select compatible cards for migration.
[0031] Through the above technical solution, this embodiment comprehensively determines whether a graphics card needs task migration and initially filters out a set of qualified migration targets, using a model. Calculate the overall state coefficient of the i-th graphics card, where the harmonic average of the two is calculated, which simultaneously considers the stability of the graphics card under high pressure and its basic health during idle time. Compared with the first preset threshold set by the system If a comparison is made, Less than If the i-th graphics card is in an unsatisfactory working state, its tasks need to be migrated to other graphics cards. The overall state coefficients of the n graphics cards are then calculated sequentially, and these coefficients are compared with the second preset threshold set by the system. In comparison, among them, < Filter out graphics cards whose overall performance coefficient is greater than or equal to the second preset threshold. All m graphics cards are migration objects.
[0032] As a further description of the technical solution of the present invention, the method for further selecting an adapter card for migration from m graphics cards includes: Obtain the peak computing power, memory bandwidth, and memory capacity of the j-th graphics card, where j belongs to m. Construct a mathematical model for the ji migration index coefficient, with the expression: ; In the formula, , These are the peak computing power of the j-th and i-th graphics cards, respectively. , These are the video memory bandwidths of the j-th and i-th graphics cards, respectively. , Let be the video memory capacities of the j-th and i-th graphics cards, respectively. Let be the migration index coefficient of the j-th graphics card relative to the i-th graphics card; like If the value is approximately equal to 1, it means the two cards have similar performance and can be used interchangeably; if... If the value is greater than 1, it means that j is significantly stronger than i; if If the value is much greater than 1, it means that j is significantly stronger than i, and the performance may be excessive. Calculate the migration index coefficients of m graphics cards relative to the i-th graphics card in sequence, and select the graphics card with a migration index coefficient greater than 1 and closest to 1 as the i-th graphics card task migration object.
[0033] Through the above technical solution, this embodiment provides a method for selecting the best migration target. It constructs a mathematical model of migration index coefficients by comprehensively considering three key hardware performance indicators using a geometric mean approach. This ensures that the selected target cards have no significant shortcomings in the three core hardware indicators of computing power, bandwidth, and capacity. The migration index coefficient of m graphics cards relative to the i-th graphics card is calculated in turn. The graphics card with a migration index coefficient greater than 1 and closest to 1 is selected as the i-th graphics card task migration target. This helps to maintain the load of all graphics cards at a relatively balanced level, avoids the situation where a few strong cards are overloaded and most weak cards are idle, and thus achieves true load balancing.
[0034] As a further description of the technical solution of the present invention, the status feedback module transmits the adjusted graphics card running status back to the central scheduling module in real time, dynamically adjusts the scheduling strategy, and ensures the stability of the balancing effect.
[0035] It should be noted that the formulas in this application are all dimensionless numerical calculations, implemented using existing technology, and need no further explanation. The formulas are derived from software simulations using a large amount of collected data, and are the closest to the real situation. The thresholds, threshold ranges, and coefficients involved in this application are all empirical values, and their selection is made by those skilled in the art based on the actual situation.
[0036] The foregoing has provided a detailed description of one embodiment of the present invention, but this description is merely a preferred embodiment and should not be construed as limiting the scope of the invention. All equivalent variations and modifications made within the scope of the claims of this invention should still fall within the patent coverage of this invention.
Claims
1. A server multi-GPU load balancing management system, characterized in that, The system includes a load acquisition module, a central scheduling module, a graphics card control module, and a status feedback module. The load acquisition module is used to establish data communication connections with each graphics card in the server and collect the load parameters of each graphics card in real time. The load parameters include core temperature, memory usage, computing power utilization and task queue length. The central scheduling module is communicatively connected to the load acquisition module, the graphics card control module, and the status feedback module, respectively. It is used to receive the load parameters transmitted by the load acquisition module, perform comprehensive analysis on the load parameters based on a preset load balancing algorithm, and generate task scheduling instructions. The graphics card control module is communicatively connected to the central scheduling module and each graphics card, and is used to receive task scheduling instructions sent by the central scheduling module, and to perform task allocation and migration according to the task scheduling instructions; The status feedback module is communicatively connected to the graphics card control module and each graphics card. It is used to collect the real-time operating status parameters of each graphics card after load adjustment and to send the real-time operating status parameters back to the central scheduling module to form a load balancing closed-loop control.
2. The server multi-GPU load balancing management system according to claim 1, characterized in that, The load acquisition module includes a parameter acquisition unit and a data preprocessing unit; The working process of the parameter acquisition unit includes: numbering all the graphics cards on the server, with the numbers being 1, 2, ..., n in sequence; Collect the temperature, memory usage, computing power utilization, and task queue length of each of the n graphics card cores. The data preprocessing unit's operation includes: noise reduction, outlier removal, and standardization of the collected load parameters, wherein outlier removal employs a 3D algorithm. The standardization process uses a normalization method to convert load parameters to the same standard.
3. The server multi-GPU load balancing management system according to claim 2, characterized in that, The working process of the central scheduling module includes: Obtain the average values of various load parameters for the i-th graphics card during the peak hours of the server's current work cycle. Based on the obtained average values of the parameters, construct a mathematical model for the first state coefficient of the i-th graphics card, expressed as: ; In the formula, Let be the average core temperature during the peak period of the current working cycle of the i-th graphics card. The upper limit temperature set for the i-th peak graphics card time period of the system. Let be the average video memory usage during peak hours in the current working cycle of the i-th graphics card. Let be the computing power utilization rate of the i-th graphics card during the peak period of the current working cycle. Let be the length of the task queue during peak hours in the current working cycle of the i-th graphics card. This is the first state coefficient of the graphics card.
4. A server multi-GPU load balancing management system according to claim 3, characterized in that, The operation of the central scheduling module also includes: Obtain the average values of various load parameters for the i-th graphics card during the server's off-peak hours in the current work cycle. Based on the obtained average values, construct a mathematical model for the second state coefficient of the i-th graphics card, expressed as: ; In the formula, Let be the average core temperature during the lowest point of the current operating cycle of the i-th graphics card. The upper limit temperature set for the i-th low-frequency period of the graphics card in the system. Let be the average video memory usage during the off-peak period in the current working cycle of the i-th graphics card. Let be the computing power utilization rate of the i-th graphics card during the off-peak period in the current working cycle. Let be the length of the task queue during the off-peak period in the current working cycle of the i-th graphics card. This is the second state coefficient of the graphics card.
5. A server multi-GPU load balancing management system according to claim 4, characterized in that, The operation of the central scheduling module also includes: Based on the first and second state coefficients of the i-th graphics card, a mathematical model of the comprehensive state coefficients of the i-th graphics card is constructed, and its expression is: ; In the formula, Let i be the overall state coefficient of the i-th graphics card. Compared with the first preset threshold set by the system If a comparison is made, Less than If the i-th graphics card is in an unsatisfactory working state, then the tasks of the i-th graphics card need to be transferred to other graphics cards.
6. A server multi-GPU load balancing management system according to claim 3, characterized in that, The method for migrating the workload of the i-th graphics card to other graphics cards includes: Calculate the overall status coefficients of n graphics cards sequentially, and then compare the overall status coefficients of the n tabs with the second preset threshold set by the system. In comparison, among them, < Filter out graphics cards whose overall performance coefficient is greater than or equal to the second preset threshold. All m graphics cards are migration objects, where m belongs to n; From the m graphics cards, further select compatible cards for migration.
7. A server multi-GPU load balancing management system according to claim 6, characterized in that, The method for further selecting compatible graphics cards for migration from among m graphics cards includes: Obtain the peak computing power, memory bandwidth, and memory capacity of the j-th graphics card, where j belongs to m. Construct a mathematical model for the ji migration index coefficient, with the expression: ; In the formula, , These are the peak computing power of the j-th and i-th graphics cards, respectively. , These are the video memory bandwidths of the j-th and i-th graphics cards, respectively. , Let be the video memory capacities of the j-th and i-th graphics cards, respectively. Let be the migration index coefficient of the j-th graphics card relative to the i-th graphics card. , and These are the weighting coefficients; like If the value is approximately equal to 1, it means the two cards have similar performance and can be used interchangeably; if... If the value is greater than 1, it means that j is significantly stronger than i; if If the value is much greater than 1, it means that j is significantly stronger than i, and the performance may be excessive. Calculate the migration index coefficients of m graphics cards relative to the i-th graphics card in sequence, and select the graphics card with a migration index coefficient greater than 1 and closest to 1 as the i-th graphics card task migration object.
8. A server multi-GPU load balancing management system according to claim 7, characterized in that, The status feedback module transmits the adjusted graphics card operating status back to the central scheduling module in real time, dynamically adjusts the scheduling strategy, and ensures the stability of the balancing effect.
Citation Information
Patent Citations
Machine room energy consumption and computing power balance optimization method and system based on swarm intelligence
CN120386612A
Server GPU (Graphics Processing Unit) computing power distribution method and system and server
CN120743469A