Machine room computing power resource unified scheduling monitoring system based on dynamic load balancing

By constructing a unified scheduling and monitoring system for data center computing resources with dynamic load balancing, the problem of dynamic adjustment of data center computing resource scheduling was solved, achieving efficient resource utilization and stable business operation, and providing a scientific resource allocation solution.

CN122173278APending Publication Date: 2026-06-09BEIJING HONGRUN ZHONGHE TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING HONGRUN ZHONGHE TECHNOLOGY CO LTD
Filing Date
2026-03-03
Publication Date
2026-06-09

Smart Images

  • Figure CN122173278A_ABST
    Figure CN122173278A_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of dynamic load balancing. The present application relates to a unified scheduling and monitoring system for computer room computing power resources based on dynamic load balancing. It includes a computing power node setting module, a running prediction analysis module, a demand resource updating module, a demand resource analysis module, a standby resource setting module and a resource scheduling module. The computing power node setting module is used to collect computer room computing power resources, which are divided into to-be-allocated computing power resources and used computing power resources. The present application builds a bounded computing power node state prediction system, taking the latest running state of the computing power node as the prediction center, setting the upper and lower limit boundaries of the prediction in combination with the total available computing power of the computing power resource pool, avoiding the invalid results caused by boundaryless prediction, and based on the change trend and fluctuation law of the historical running state sequence, the prediction result is more in line with the actual running characteristics of the computing power node, providing a reliable state basis for the analysis of subsequent demand computing power resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of dynamic load balancing technology, and more specifically, to a unified scheduling and monitoring system for data center computing resources based on dynamic load balancing. Background Technology

[0002] In the context of digital and intelligent development, data centers, as the core carriers of computing power output, play a crucial role in ensuring stable business operations and improving resource utilization through the scheduling and monitoring of their internal hardware computing resources, such as CPUs, GPUs, memory, storage, and network bandwidth.

[0003] Currently, load balancing scheduling is typically carried out based on fixed thresholds, static rules, or simple round-robin and weighted algorithms. Some improved solutions combine historical operating data for simple load trend judgments, and scheduling decisions are mostly based on the resource status of single nodes or local clusters. This results in the inability to dynamically adjust thresholds according to the actual resource usage of computing nodes, which easily leads to problems such as high-load nodes having excessively low thresholds that frequently trigger scheduling, and low-load nodes having excessively high thresholds that waste resources. The adaptability and flexibility of scheduling are insufficient. At the same time, the prediction boundary is not set in conjunction with the total available computing power of the overall computing power resource pool in the data center, which easily leads to invalid situations where the prediction results exceed the physical resource carrying capacity. Furthermore, there is a lack of confidence analysis of the prediction results. Scheduling decisions based on unreliable prediction results can easily cause business operation failures. In order to reduce these situations, a unified scheduling and monitoring system for data center computing power resources based on dynamic load balancing is proposed. Summary of the Invention

[0004] The purpose of this invention is to provide a unified scheduling and monitoring system for data center computing resources based on dynamic load balancing, so as to solve the problems mentioned in the background art.

[0005] To achieve the above objectives, a unified scheduling and monitoring system for data center computing resources based on dynamic load balancing is provided, including a computing node setting module, an operation prediction and analysis module, a demand resource update module, a demand resource analysis module, a reserve resource setting module, and a resource scheduling module. The computing power node setting module is used to collect computing power resources in the computer room, divide them into computing power resources to be allocated and computing power resources already used, aggregate the computing power resources to be allocated to form a computing power resource pool, and set computing power nodes according to the service objects of the computing power resources already used. The operation prediction and analysis module is used to obtain the historical operation status of each computing power node, extract the latest operation status as the prediction center, set the prediction boundary according to the total computing power resources of the computing power resource pool, and perform prediction analysis by combining the prediction center, prediction boundary and historical operation status to obtain a list of predicted operation statuses for each computing power node. The demand resource update module is used to analyze the demand computing resources under the corresponding predicted operating status based on the used computing resources and the predicted operating status list of each computing node, dynamically set the balancing threshold based on the used computing resources, and update the demand computing resources according to the balancing threshold. The demand resource analysis module is used to perform confidence analysis on the list of predicted operating states based on historical operating states to obtain the confidence of each predicted operating state, and to perform resource correlation analysis on all demand computing resources to obtain the correlation of each demand computing resource. The reserved resource setting module is used to extract the required computing power resources with the largest difference value among each computing power node, set a confidence level requirement threshold, filter out the required computing power resources with a confidence level lower than the threshold, and set the number of reserved computing power resources for each computing power node according to the confidence level of the required computing power resources with the largest difference value. The resource scheduling module is used to select the nearest available computing resources for scheduling when the running status of a computing node is updated; for computing nodes that need to be scaled down, some of the used computing resources are migrated back to the computing resource pool.

[0006] As a further improvement to this technical solution, the computing power node setting module uniformly collects and counts hardware computing power resources such as CPU, GPU, memory, storage, and network bandwidth in the computer room, classifies hardware computing power resources in an idle and allocable state as computing power resources to be allocated, and classifies hardware computing power resources in a service-carrying and occupied state as computing power resources already used. At the same time, all unallocated computing resources will be uniformly aggregated and standardized to form a computing resource pool. Based on the business clusters supported by the used computing resources, the corresponding used computing resources are divided into independently managed computing nodes.

[0007] As a further improvement to this technical solution, the operation prediction and analysis module collects historical operation status of computing nodes and sorts the collected historical operation status according to the occurrence time to form a continuous historical operation status sequence. Extract the latest occurrence time of the operation status data from the historical operation status sequence as the latest operation status; The latest operating status is set as the prediction baseline center, and the total available computing power of the computing power resource pool is used as the upper limit boundary of the prediction, and zero computing power occupancy is used as the lower limit boundary of the prediction to construct the prediction interval. Based on the changing trends and fluctuation patterns of historical operating state sequences, time-series extrapolation calculations are performed within the prediction interval to obtain the predicted operating states at multiple future moments, which are then combined to form a list of predicted operating states.

[0008] As a further improvement to this technical solution, the required resource update module collects the predicted operating status list of each computing power node in a unified manner and establishes a correspondence between computing power nodes and predicted operating status. The required computing resources are analyzed by combining the used computing resources of each computing node with the predicted operating status of its corresponding predicted operating status list. For each predicted operating status, the total amount of computing resources required for that predicted operating status is obtained by combining the current usage of computing resources with the predicted load change ratio, and this amount is taken as the required computing resources.

[0009] As a further improvement to this technical solution, in the demand resource update module, a balance threshold is dynamically set based on the used computing power resources, and then the demand computing power resources are updated according to the balance threshold to obtain new demand computing power resources. The new demand computing power resources are calculated in both the demand resource analysis module and the pre-set resource module. Among them, the equilibrium threshold is positively correlated with the computing power resources already used; The higher the amount of computing power already used, the higher the equilibrium threshold. The lower the amount of computing power used, the lower the equilibrium threshold. The required computing power resources are compared with the equilibrium threshold. When the required computing power resources exceed the equilibrium threshold, the required computing power resources are expanded and updated. Conversely, when the required computing power resources are lower than the equilibrium threshold, the required computing power resources are scaled down and updated.

[0010] As a further improvement to this technical solution, in the demand resource analysis module, a confidence analysis is performed on the list of predicted operating states based on the historical operating states of each computing power node. Based on the degree of fit between the historical operating states and the predicted operating states, and the historical fluctuation deviation, the credibility probability of each predicted operating state is calculated as the confidence level. By combining the required computing power resources corresponding to all predicted operating states, a resource correlation analysis is performed. By traversing all required computing power resources of all computing power nodes, the resource quantity difference between the target required computing power resources and other required computing power resources is calculated, thereby obtaining the correlation degree of each predicted operating state. The smaller the difference in the amount of computing power required by the target and other required computing power, the higher the correlation. The greater the difference in the amount of computing power required by the target user compared to other users, the lower the correlation.

[0011] As a further improvement to this technical solution, in the pre-resource setting module, in the set of required computing resources of the same computing power node, the difference between each required computing resource and the currently used computing resources is calculated, and the required computing resource with the largest difference is determined as the required computing resource with the largest difference value. Set a confidence threshold and compare the confidence level of each required computing power resource with the confidence threshold. When the confidence level is higher than the required confidence level threshold, the computing resources required for that demand will be reserved. Conversely, if the confidence level is lower than the required threshold, the computing power resources required will be filtered out.

[0012] As a further improvement to this technical solution, in the reserve resource setting module, when the confidence level of the required computing power resource with the largest difference value is higher than the confidence level requirement threshold, the required computing power resource is used as the only reserve computing power resource for the computing power node; when the confidence level of the required computing power resource with the largest difference value is lower than the confidence level requirement threshold, the required computing power resource is used as a fixed reserve computing power resource for the computing power node, and the required computing power resource with the highest correlation is selected from the reserved required computing power resources as a supplementary reserve computing power resource, so that the computing power node has two reserve computing power resources.

[0013] As a further improvement to this technical solution, in the resource scheduling module, when the running status of the computing power node changes, the distance between the currently used computing power resources and each reserve computing power resource is calculated, and the nearest reserve computing power resource is selected for matching scheduling. When the scheduling result is to reduce capacity, the used computing resources in the computing power nodes that exceed the reserved computing power resources will be migrated and released and recycled to the computing power resource pool. When the scheduling result is expansion, the reserved computing resources are allocated from the computing resource pool to the corresponding computing nodes; Meanwhile, when the computing power node remains unchanged, but the distance between the computing power resources corresponding to the changed computing power node and the unchanged computing power node is less than the reserved computing power resources of the unchanged computing power node, the reserved computing power resources of the unchanged computing power node are scheduled to the changed computing power node, and then the reserved computing power resources corresponding to the unchanged computing power node and the changed computing power node are updated synchronously.

[0014] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. In this unified scheduling and monitoring system for data center computing resources based on dynamic load balancing, a bounded computing node status prediction system is constructed. The latest operating status of the computing node is used as the prediction center, and the upper and lower limits of the prediction are set in combination with the total available computing power of the computing resource pool. This avoids the problem of invalid results caused by unbounded prediction. At the same time, time-series extrapolation is performed based on the changing trends and fluctuation patterns of historical operating status sequences, so that the prediction results are more in line with the actual operating characteristics of the computing nodes, providing a reliable status basis for the subsequent analysis of computing resource demand.

[0015] 2. In this unified scheduling and monitoring system for data center computing resources based on dynamic load balancing, the required computing resources are accurately calculated and dynamically calibrated through the required resource update module. By combining the used computing resources with the predicted load change ratio, the required computing resources are calculated, ensuring the accuracy of the demand calculation. Furthermore, the system dynamically sets the balancing threshold based on the actual situation of the used computing resources, achieving a positive correlation between the balancing threshold and the load status of the computing nodes. Through threshold comparison, the system completes the expansion or contraction update of the required computing resources, making the required computing resources more in line with the actual resource carrying capacity and business operation needs of the data center, and providing accurate resource demand references for subsequent scheduling decisions.

[0016] 3. In this unified scheduling and monitoring system for data center computing resources based on dynamic load balancing, the system achieves a quantitative assessment of the predicted operating status and corresponding computing resource requirements through a dual-dimensional analysis of confidence and correlation. The confidence level is calculated based on the degree of fit between historical and predicted operating statuses and historical fluctuation deviations, effectively screening out reliable computing resources in need and avoiding scheduling decisions based on ineffective predictions. Furthermore, the correlation level is obtained by calculating the quantitative differences between required computing resources, providing a scientific basis for the subsequent supplementary allocation of reserve computing resources and making the allocation of reserve computing resources more rational. Attached Figure Description

[0017] Figure 1 This is a schematic diagram of the structure of the unified scheduling and monitoring system for data center computing resources based on dynamic load balancing, as described in this invention. Figure 2 This is a flowchart illustrating the computing node setting module of the present invention; Figure 3 This is a flowchart illustrating the operation of the predictive analysis module of the present invention; Figure 4 This is a flowchart illustrating the resource update module of the present invention. Figure 5 This is a flowchart illustrating the resource demand analysis module of the present invention. Figure 6 A flowchart illustrating the pre-resource setting module of this invention; Figure 7 This is a flowchart illustrating the resource scheduling module of the present invention. Detailed Implementation

[0018] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0019] Please see Figures 1-7 As shown, the purpose of this embodiment is to provide a unified scheduling and monitoring system for data center computing resources based on dynamic load balancing, including a computing node setting module, an operation prediction and analysis module, a demand resource update module, a demand resource analysis module, a reserve resource setting module, and a resource scheduling module; The computing power node setting module is used to collect computing power resources in the computer room, divide them into computing power resources to be allocated and computing power resources already used, aggregate the computing power resources to be allocated to form a computing power resource pool, and set computing power nodes according to the service objects of the computing power resources already used. The computing node setting module collects and statistically analyzes hardware computing resources such as CPU, GPU, memory, storage, and network bandwidth in the computer room. The collected data is time-aligned, unit-normalized, and outlier-filtered to form resource status data in a unified format; Hardware computing resources that are idle and available for allocation are classified as computing resources to be allocated, and hardware computing resources that are occupied by services are classified as computing resources already in use. At the same time, all unallocated computing resources will be uniformly aggregated and standardized to form a computing resource pool. Based on the business clusters supported by the used computing resources, the corresponding used computing resources are divided into independently managed computing nodes.

[0020] It iterates through all used computing resources, identifies the business type and business cluster identifier they support, and then aggregates the used computing resources belonging to the same business cluster, creating an independent management unit for each business cluster, which is defined as a computing node.

[0021] The predictive analysis module is used to obtain the historical operating status of each computing power node, extract the latest operating status as the prediction center, set the prediction boundary according to the total computing power resources of the computing power resource pool, and perform predictive analysis by combining the prediction center, prediction boundary and historical operating status to obtain a list of predicted operating status of each computing power node. In the predictive analysis module, the historical operating status of the computing nodes is collected, and the collected historical operating status is sorted according to the occurrence time to form a continuous historical operating status sequence. Extract the latest occurrence time of the operation status data from the historical operation status sequence as the latest operation status; The operating status indicators of computing nodes are collected at fixed time intervals, including computing load and resource utilization. The collected historical operating statuses are arranged in chronological order to form a continuous time series. Then, the latest data in the historical operating status series is selected as the latest operating status. The latest operating status is set as the prediction baseline center, and the total available computing power of the computing power resource pool is used as the upper limit boundary of the prediction, and zero computing power occupancy is used as the lower limit boundary of the prediction to construct the prediction interval. Based on the changing trends and fluctuation patterns of historical operating state sequences, time-series extrapolation calculations are performed within the prediction interval to obtain predicted operating states for multiple future moments. These predicted operating states are then combined to form a list of predicted operating states, as shown in the following formula: ; ; in, The predicted running state at the k-th time in the future. This is a sequence of historical operating states. For time series prediction functions, As the prediction benchmark center, To predict the lower limit, This represents the upper limit of the prediction.

[0022] The demand resource update module is used to analyze the demand computing resources under the corresponding predicted operating status based on the list of used computing resources and predicted operating status of each computing node, dynamically set the balancing threshold based on the used computing resources, and update the demand computing resources according to the balancing threshold. In the resource demand update module, the predicted operating status list of each computing power node is collected in a unified manner, and the correspondence between computing power nodes and predicted operating status is established. Traverse all computing power nodes, collect the prediction running status list corresponding to each node, use the computing power node ID as an index to bind the node information with the corresponding prediction running status list, establish a one-to-one mapping relationship, and then generate a collection table.

[0023] The required computing resources are analyzed by combining the used computing resources of each computing node with the predicted operating status of its corresponding predicted operating status list. For each predicted operating status, the total amount of computing resources required for that predicted operating status is obtained by combining the current usage of computing resources with the predicted load change ratio, and this amount is taken as the required computing resources.

[0024] Extract the current used computing resources of each computing node, covering core resources such as CPU, GPU, and memory. Then extract the load data corresponding to each predicted running status from the list of predicted running statuses of each node, and calculate the load change ratio of the predicted status relative to the current used resources. For each predicted operating state of each computing node, substitute the used resource usage and load change ratio to calculate the total computing resources required in that state. Verify the calculation results to ensure that the required computing resources are not less than 0 and do not exceed the total available computing power of the computing resource pool. At the same time, determine that the calculation result is the required computing resources in the corresponding predicted operating state and bind and archive it with the node and the predicted state. In the demand resource update module, the balancing threshold is dynamically set based on the used computing power resources. Then, the demand computing power resources are updated according to the balancing threshold to obtain new demand computing power resources. The new demand computing power resources are calculated in both the demand resource analysis module and the pre-set resource module. Extract the current usage of computing resources for each computing node, normalize it to the usage ratio in the range of 0-1, and then calculate the balance threshold for the node based on the positive correlation rule that the higher the usage of computing power, the higher the balance threshold. Ensure that the balance threshold is not lower than the minimum available unit of the computing resource pool and does not exceed the total available computing power of the computing resource pool. Among them, the equilibrium threshold is positively correlated with the computing power resources already used; The higher the amount of computing power already used, the higher the equilibrium threshold. The lower the used computing resources, the lower the equilibrium threshold, as shown in the formula below: ; in, Let be the equilibrium threshold for the i-th computing node. This is the threshold adjustment coefficient. This represents the normalized percentage of computing power used by the i-th computing node. The total available computing power in the computing power resource pool; The required computing power resources are compared with the equilibrium threshold. When the required computing power resources exceed the equilibrium threshold, the required computing power resources are expanded and updated. Conversely, when the required computing power resources are lower than the equilibrium threshold, the required computing power resources are scaled down and updated.

[0025] The demand resource analysis module is used to perform confidence analysis on the list of predicted operating states based on historical operating states to obtain the confidence of each predicted operating state, and to perform resource correlation analysis on all demand computing resources to obtain the correlation of each demand computing resource. In the demand resource analysis module, based on the historical operating status of each computing power node, a confidence analysis is performed on the list of predicted operating statuses. Based on the degree of fit between the historical operating status and the predicted operating status, and the historical fluctuation deviation, the confidence probability of each predicted operating status is calculated as the confidence level. By calculating the fitting error (such as mean squared error) between historical and predicted states, the degree of fit between the predicted value and historical patterns is reflected. Simultaneously, the standard deviation of the historical state sequence is calculated to reflect the stability of node load (smaller fluctuations indicate higher prediction reliability). Combining the fit degree and fluctuation deviation, the result is normalized to a confidence probability in the 0-1 interval, which is used as the confidence level of the predicted operating state. The formula is as follows: ; in, Let be the standard deviation (fluctuation deviation) of the historical operating status of the i-th computing node. For the number of moments, This represents the historical running state of the i-th node at time p. This represents the average historical running status of the i-th node; ; in, Let be the mean square error (fitness) of the predicted state of the i-th node at time k. This represents the predicted running state of the i-th node at time k in the future. ; in, Let be the confidence level of the predicted state of the i-th node at time k. and These are the weighting coefficients; By combining the required computing power resources corresponding to all predicted operating states, a resource correlation analysis is performed. By traversing all required computing power resources of all computing power nodes, the resource quantity difference between the target required computing power resources and other required computing power resources is calculated, thereby obtaining the correlation degree of each predicted operating state. The required computing resources of all computing power nodes and all predicted running states are aggregated to form a global set of required resources. Then, each target required computing power resource in the set is traversed, and the absolute difference between it and all other required computing power resources in the set is calculated. The difference is then mapped inversely to the correlation degree in the range of 0-1 (the smaller the difference, the closer the correlation degree is to 1; the larger the difference, the closer the correlation degree is to 0). The smaller the difference in the amount of computing power required by the target and other required computing power, the higher the correlation. The greater the difference in the amount of computing power required by the target user compared to other users, the lower the correlation.

[0026] The reserve resource setting module is used to extract the required computing power resources with the largest difference value among each computing power node, set a confidence level threshold, filter out required computing power resources with a confidence level lower than the threshold, and set the number of reserve computing power resources for each computing power node based on the confidence level of the required computing power resources with the largest difference value. In the pre-resource setting module, in the set of required computing resources for the same computing power node, the difference between each required computing resource and the currently used computing resources is calculated, and the required computing resource with the largest difference is determined as the required computing resource with the largest difference. Set a confidence threshold (0.7, which can be adjusted according to the scenario), and compare the confidence level corresponding to each required computing power resource with the confidence threshold. When the confidence level is higher than the required confidence level threshold, the computing resources required for that demand will be reserved. Conversely, if the confidence level is lower than the required threshold, the computing power resources required will be filtered out.

[0027] In the reserve resource setting module, when the confidence level of the required computing power resource with the largest difference value is higher than the confidence level requirement threshold, the required computing power resource is used as the only reserve computing power resource for the computing power node; when the confidence level of the required computing power resource with the largest difference value is lower than the confidence level requirement threshold, the required computing power resource is used as the fixed reserve computing power resource for the computing power node, and the required computing power resource with the highest correlation is selected from the reserved required computing power resources as a supplementary reserve computing power resource, so that the computing power node has two reserve computing power resources.

[0028] The resource scheduling module is used to select the nearest available computing resources for scheduling when the running status of computing nodes is updated; for computing nodes that need to be scaled down, some of the used computing resources are migrated back to the computing resource pool.

[0029] In the resource scheduling module, when the running status of a computing node changes, the distance (numerical difference) between the currently used computing resources and each reserve computing resource is calculated, and the nearest reserve computing resource is selected for matching scheduling. When the scheduling result is to reduce capacity, the used computing resources in the computing power nodes that exceed the reserved computing power resources will be migrated and released and recycled to the computing power resource pool. When the scheduling result is expansion, the reserved computing resources are allocated from the computing resource pool to the corresponding computing nodes; Meanwhile, when the computing power node remains unchanged, but the distance between the computing power resources corresponding to the changed computing power node and the unchanged computing power node is less than the reserved computing power resources of the unchanged computing power node, the reserved computing power resources of the unchanged computing power node are scheduled to the changed computing power node, and then the reserved computing power resources corresponding to the unchanged computing power node and the changed computing power node are updated synchronously.

[0030] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely preferred examples and are not intended to limit the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of the present invention is defined by the appended claims and their equivalents.

Claims

1. A unified scheduling and monitoring system for data center computing resources based on dynamic load balancing, characterized in that: It includes a computing node setting module, an operation prediction and analysis module, a demand resource update module, a demand resource analysis module, a reserve resource setting module, and a resource scheduling module; The computing power node setting module is used to collect computing power resources in the computer room, divide them into computing power resources to be allocated and computing power resources already used, aggregate the computing power resources to be allocated to form a computing power resource pool, and set computing power nodes according to the service objects of the computing power resources already used. The operation prediction and analysis module is used to obtain the historical operation status of each computing power node, extract the latest operation status as the prediction center, set the prediction boundary according to the total computing power resources of the computing power resource pool, and perform prediction analysis by combining the prediction center, prediction boundary and historical operation status to obtain a list of predicted operation statuses for each computing power node. The demand resource update module is used to analyze the demand computing resources under the corresponding predicted operating status based on the used computing resources and the predicted operating status list of each computing node, dynamically set the balancing threshold based on the used computing resources, and update the demand computing resources according to the balancing threshold. The demand resource analysis module is used to perform confidence analysis on the list of predicted operating states based on historical operating states to obtain the confidence of each predicted operating state, and to perform resource correlation analysis on all demand computing resources to obtain the correlation of each demand computing resource. The reserved resource setting module is used to extract the required computing power resources with the largest difference value among each computing power node, set a confidence level requirement threshold, filter out the required computing power resources with a confidence level lower than the threshold, and set the number of reserved computing power resources for each computing power node according to the confidence level of the required computing power resources with the largest difference value. The resource scheduling module is used to select the nearest available computing resources for scheduling when the running status of a computing node is updated; for computing nodes that need to be scaled down, some of the used computing resources are migrated back to the computing resource pool.

2. The unified scheduling and monitoring system for data center computing resources based on dynamic load balancing according to claim 1, characterized in that: In the computing power node setting module, the hardware computing power resources such as CPU, GPU, memory, storage and network bandwidth in the computer room are uniformly collected and statistically analyzed. Hardware computing power resources in an idle and allocable state are classified as computing power resources to be allocated, and hardware computing power resources in a service-carrying and occupied state are classified as computing power resources already used. At the same time, all unallocated computing resources will be uniformly aggregated and standardized to form a computing resource pool. Based on the business clusters supported by the used computing resources, the corresponding used computing resources are divided into independently managed computing nodes.

3. The unified scheduling and monitoring system for data center computing resources based on dynamic load balancing according to claim 1, characterized in that: In the operation prediction and analysis module, the historical operation status of the computing power nodes is collected, and the collected historical operation status is sorted according to the occurrence time to form a continuous historical operation status sequence. Extract the latest occurrence time of the operation status data from the historical operation status sequence as the latest operation status; The latest operating status is set as the prediction baseline center, and the total available computing power of the computing power resource pool is used as the upper limit boundary of the prediction, and zero computing power occupancy is used as the lower limit boundary of the prediction to construct the prediction interval. Based on the changing trends and fluctuation patterns of historical operating state sequences, time-series extrapolation calculations are performed within the prediction interval to obtain the predicted operating states at multiple future moments, which are then combined to form a list of predicted operating states.

4. The unified scheduling and monitoring system for data center computing resources based on dynamic load balancing according to claim 1, characterized in that: In the resource demand update module, the predicted operating status list of each computing power node is collected in a unified manner, and a correspondence between computing power nodes and predicted operating status is established. The required computing resources are analyzed by combining the used computing resources of each computing node with the predicted operating status of its corresponding predicted operating status list. For each predicted operating status, the total amount of computing resources required for that predicted operating status is obtained by combining the current usage of computing resources with the predicted load change ratio, and this amount is taken as the required computing resources.

5. The unified scheduling and monitoring system for data center computing resources based on dynamic load balancing according to claim 1, characterized in that: In the demand resource update module, a balance threshold is dynamically set based on the used computing power resources, and then the demand computing power resources are updated according to the balance threshold to obtain new demand computing power resources. The new demand computing power resources are calculated in both the demand resource analysis module and the pre-set resource module. Among them, the equilibrium threshold is positively correlated with the computing power resources already used; The higher the amount of computing power already used, the higher the equilibrium threshold. The lower the amount of computing power used, the lower the equilibrium threshold. The required computing power resources are compared with the equilibrium threshold. When the required computing power resources exceed the equilibrium threshold, the required computing power resources are expanded and updated. Conversely, when the required computing power resources are lower than the equilibrium threshold, the required computing power resources are scaled down and updated.

6. The unified scheduling and monitoring system for data center computing resources based on dynamic load balancing according to claim 1, characterized in that: In the demand resource analysis module, a confidence analysis is performed on the list of predicted operating states based on the historical operating states of each computing power node. Based on the degree of fit between the historical operating states and the predicted operating states, and the historical fluctuation deviation, the confidence probability of each predicted operating state is calculated as the confidence level. By combining the required computing power resources corresponding to all predicted operating states, a resource correlation analysis is performed. By traversing all required computing power resources of all computing power nodes, the resource quantity difference between the target required computing power resources and other required computing power resources is calculated, thereby obtaining the correlation degree of each predicted operating state. The smaller the difference in the amount of computing power required by the target and other required computing power, the higher the correlation. The greater the difference in the amount of computing power required by the target user compared to other users, the lower the correlation.

7. The unified scheduling and monitoring system for data center computing resources based on dynamic load balancing according to claim 1, characterized in that: In the pre-resource setting module, in the set of required computing resources of the same computing power node, the difference between each required computing resource and the currently used computing resources is calculated, and the required computing resource with the largest difference is determined as the required computing resource with the largest difference. Set a confidence threshold and compare the confidence level of each required computing power resource with the confidence threshold. When the confidence level is higher than the required confidence level threshold, the computing resources required for that demand will be reserved. Conversely, if the confidence level is lower than the required threshold, the computing power resources required will be filtered out.

8. The unified scheduling and monitoring system for data center computing resources based on dynamic load balancing according to claim 1, characterized in that: In the reserved resource setting module, when the confidence level of the required computing power resource with the largest difference value is higher than the confidence level requirement threshold, the required computing power resource is used as the only reserved computing power resource for the computing power node; when the confidence level of the required computing power resource with the largest difference value is lower than the confidence level requirement threshold, the required computing power resource is used as a fixed reserved computing power resource for the computing power node, and the required computing power resource with the highest correlation is selected from the reserved required computing power resources as a supplementary reserved computing power resource, so that the computing power node has two reserved computing power resources.

9. The unified scheduling and monitoring system for data center computing resources based on dynamic load balancing according to claim 1, characterized in that: In the resource scheduling module, when the running status of the computing power node changes, the distance between the currently used computing power resources and each reserve computing power resource is calculated, and the nearest reserve computing power resource is selected for matching scheduling. When the scheduling result is to reduce capacity, the used computing resources in the computing power nodes that exceed the reserved computing power resources will be migrated and released and recycled to the computing power resource pool. When the scheduling result is expansion, the reserved computing resources are allocated from the computing resource pool to the corresponding computing nodes; Meanwhile, when the computing power node remains unchanged, but the distance between the computing power resources corresponding to the changed computing power node and the unchanged computing power node is less than the reserved computing power resources of the unchanged computing power node, the reserved computing power resources of the unchanged computing power node are scheduled to the changed computing power node, and then the reserved computing power resources corresponding to the unchanged computing power node and the changed computing power node are updated synchronously.