A data center cluster distributed cooperative regulation method and a regulation system
By establishing a power-computing power coordination optimization model in a data center cluster and employing a distributed iterative algorithm based on the alternating direction multiplier method, the problem of independent scheduling of power and computing power was solved. This achieved joint optimal resource allocation and privacy protection, reduced operating costs, and improved system reliability and renewable energy absorption capacity.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- STATE GRID ANHUI ELECTRIC POWER CO LTD
- Filing Date
- 2026-02-06
- Publication Date
- 2026-06-05
AI Technical Summary
In the existing data center cluster scheduling model, power scheduling and computing power allocation are independent of each other, resulting in low resource utilization efficiency, high operating costs, and the centralized scheduling method has the risk of data privacy leakage and single point of failure.
A collaborative optimization model for data center clusters with the goal of minimizing total system cost is established. A power-computing power coordinated optimization sub-model is constructed, and a distributed iterative algorithm using the alternating direction multiplier method is adopted to achieve the joint optimal allocation of power and computing power resources. Furthermore, a distributed iterative solution algorithm is used to protect commercial privacy.
It achieves optimal joint allocation of power and computing resources, reduces the total operating cost of data center clusters, protects business privacy, has good convergence and scalability, provides flexible load regulation capabilities, and supports high-proportion renewable energy consumption.
Smart Images

Figure CN122152504A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of information technology and energy management, and in particular to a distributed collaborative control method and control system for data center clusters. Background Technology
[0002] Traditional centralized cluster collaborative control methods rely on a central controller to collect full operational information from each data center for unified decision-making. While this method can obtain a globally optimal solution, it can leak data center trade secrets in practical applications and suffers from heavy communication and computational burdens, posing a single point of failure risk. To address this, many studies have adopted distributed optimization methods. In distributed methods, agents do not need to upload local information to a centralized center; instead, they perform local optimization calculations and collaborate by exchanging limited information, effectively protecting the privacy of multiple agents. However, in existing data center cluster scheduling models, power scheduling and computing power allocation are often independent, failing to consider the coordination of power and computing power, resulting in low resource utilization efficiency and high operating costs.
[0003] In recent years, research has attempted to incorporate electricity costs into data center scheduling models, attracting widespread attention and research from scholars both domestically and internationally. Typical algorithms include load transfer methods based on dynamic electricity price response, green scheduling algorithms considering renewable energy, and distributed collaborative frameworks based on game theory. However, the application of such algorithms in the field of collaborative operation technology between power systems and data centers still faces three key challenges that urgently need to be addressed: First, existing research lacks sophisticated collaborative optimization models that integrate multiple cost factors. Traditional modeling methods often consider only electricity costs or computing power allocation, failing to incorporate real-time electricity prices, computing power transmission costs, and the control costs of various adjustable resources into a unified framework. This results in models that cannot accurately depict the real operation of data center clusters, and optimization results that deviate from the actual optimal. Secondly, existing methods typically treat electricity and computing power as two separate issues, lacking an interactive model that reflects the impact of computing power scheduling on electricity consumption and the impact of electricity costs on computing power flow, thus limiting the system's potential for further improvement in energy efficiency and economy. Third, centralized scheduling methods face serious risks of data privacy breaches and single points of failure. To achieve global optimization, centralized methods need to collect sensitive business secrets such as detailed internal energy consumption models, load data, and customer information from each data center. This not only poses a significant privacy risk but also introduces operational reliability risks due to its high degree of centralization. Summary of the Invention
[0004] This invention discloses a distributed collaborative control method for data center clusters, which reduces the total operating cost of data center clusters and protects business privacy.
[0005] This invention also discloses a distributed collaborative control system for data center clusters, the method comprising: Establish a collaborative optimization model for data center clusters with the goal of minimizing total system cost; Based on the aforementioned collaborative optimization model, a local power-computing power coordination optimization sub-model is constructed for each data center. Based on the aforementioned sub-model, a collaborative optimization model for inter-data center computing load balancing and transmission balancing is constructed. Based on the aforementioned collaborative optimization model, the global optimization problem is decomposed into local optimization subproblems that can be solved in parallel and global coordination subproblems. Based on the local optimization subproblem and the global coordination subproblem, a distributed iterative solution algorithm using the alternating direction multiplier method is designed. The decision to continue iteration is based on the results of the distributed iterative solution algorithm.
[0006] Furthermore, the collaborative optimization model includes minimizing the objective function and the total computing power balance constraint function: The objective function to be minimized is: ; In the formula, t is the index of the scheduling period, t=1,2,......,T; T is the total number of scheduling periods; k is the index of the data center, k=1,2,......,N; N is the number of data centers in the data center cluster; This represents the time-of-use electricity price at time t, expressed in yuan per kilowatt-hour. This represents the electrical power purchased by the k-th data center from the power grid at time t, in kilowatts. This represents the on-grid electricity price at time t, expressed in yuan per kilowatt-hour. This represents the electrical power sold to the grid by the k-th data center at time t, in kilowatts. This represents the transmission price per unit of computing power, expressed in yuan per tera-point operation. This represents the computing power transmitted from the k-th data center to the j-th data center at time t, measured in teraflops per hour. The total computing power balance constraint function is: ; In the formula, This represents the local computing load processed by the k-th data center at time t; This represents the total computing power workload required to be completed at time t.
[0007] Furthermore, the power-computing power coordinated optimization sub-model is as follows: ; In the formula, This represents the computing power transmitted from the k-th data center to the j-th data center at time t. The Lagrange multiplier representing the consistency constraint of computing power transmission at time t; This represents the transmission volume published by data center j in the previous iteration; This represents the penalty parameter of the ADMM algorithm. It is the Lagrange multiplier corresponding to the total computing power load balance constraint function, and is the system marginal cost that satisfies the unit total computing power requirement at time t. This indicates the local computing load published by data center j in the previous iteration.
[0008] Furthermore, the constraints of the objective function are as follows: Server utilization constraints: ; Constraints on the range of computing power transmission: ; Computing capacity constraints: ; Power balance constraints: ; In the formula, This represents the server utilization rate at time t; and These represent the minimum and maximum values of server utilization, respectively. This indicates the computing power capacity of the data center; This represents the maximum scaling factor for computing power transfer. This represents the local computing load at time t; The power efficiency coefficient representing the transmission of computing power; This represents the power consumption of the cooling system at time t; This represents the power consumption of the IT equipment at time t; in, ; ; ; In the formula, This represents the base power consumption of a single server in idle state. This indicates the maximum power consumption of a single server when it is fully loaded; Indicates the overall heat transfer coefficient of a data center building; This indicates the total area of the data center server room; Indicates the coefficient of performance of the refrigeration system; This represents the total computing power received from all other data centers at time t.
[0009] Furthermore, the collaborative optimization model for computing power load balancing and transmission balancing includes a computing power transmission balancing constraint function and a total computing power load decomposition constraint function: The computing power transmission balance constraint function is: ; in, ; The total computing power load decomposition constraint is: ; in, ; In the formula, This represents the total computing load of the k-th data center during time period t. This represents the computing power processed by the k-th data center during time period t.
[0010] Furthermore, the expression for the local optimization subproblem is: ; In the formula, Represents the Lagrange multiplier of the computing power transmission consistency constraint at time t during the nth update iteration; This represents the transmission volume published by data center j during the nth update iteration.
[0011] Furthermore, the global coordination sub-problem includes the computing power transmission consistency multiplier update problem and the total computing power load balancing multiplier update problem; The calculation method for the computing power transmission consistency multiplier update problem is as follows: ; In the formula, This represents the Lagrange multiplier of the consistency constraint for computing power transmission at time t during the (n+1)th update iteration; The calculation method for the total computing power load balancing multiplier update problem is as follows: ; In the formula, and These represent the Lagrange multipliers corresponding to the total computing power load balance constraint function at the (n+1)th and nth update iterations, respectively; This represents the local computing load processed by the k-th data center at time t during the (n+1)-th update iteration.
[0012] Furthermore, based on the local optimization subproblem and the global coordination subproblem, a distributed iterative solution algorithm using the alternating direction multiplier method is designed, including: The Lagrange multipliers corresponding to the transmission consistency constraint and the total computing power load balancing constraint function calculated based on the global coordination subproblem, as well as the decision variables calculated in the local optimization subproblem, are used to initialize the coordination variables and decision variables. Each data center solves its own local optimization problem based on the coordination variable and the decision variable, and updates its local decision variable accordingly. The coordination variable is updated based on the local decision variable, and the updated coordination variable is sent to each of the data centers.
[0013] Furthermore, determining whether to continue iteration based on the results of the distributed iterative solution algorithm includes: After receiving the updated local decision variables from each of the data centers, determine whether the computing power transmission consistency error and the total computing power load balancing error of each of the data centers are both less than the preset convergence threshold or whether the number of iterations has reached the maximum number of iterations. If the convergence condition is met or the number of iterations reaches the maximum number of iterations, the iteration is terminated; otherwise, the local optimization subproblem and the global coordination subproblem are solved, and the consistency error of computing power transmission and the total computing power load balancing error of each data center are re-evaluated to determine whether the iteration termination condition is met.
[0014] On the other hand, the present invention further discloses a distributed collaborative control system for data center clusters, the system comprising: The collaborative optimization model building module is used to build a collaborative optimization model for data center clusters with the goal of minimizing the total system cost. The local sub-model construction module is used to construct local power-computing power coordination optimization sub-models for each data center based on the collaborative optimization model. The collaborative optimization model construction module is used to construct a collaborative optimization model for computing load balancing and transmission balancing between data centers based on the sub-model. The problem decomposition module is used to decompose the global optimization problem into local optimization subproblems and global coordination subproblems that can be solved in parallel, based on the collaborative optimization model. The distributed iterative solution module is used to design a distributed iterative solution algorithm using the alternating direction multiplier method based on the local optimization subproblem and the global coordination subproblem. The iteration judgment module is used to determine whether to continue iterating based on the result of the distributed iterative solution algorithm.
[0015] Compared with the prior art, the present invention has at least the following technical effects: The distributed collaborative control method for data center clusters provided by this invention establishes a collaborative optimization model that integrates electricity costs and computing power transmission costs, achieving a two-way collaborative mechanism of "controlling computing with electricity and adjusting electricity with computing power," thus enabling joint optimal allocation of electricity and computing resources. Employing a distributed solution framework based on the alternating direction multiplier method, each data center only needs to exchange a small amount of boundary information such as local load decisions and computing power transmission plans with the coordinator, without disclosing sensitive data such as internal energy consumption models and customer loads, effectively protecting the commercial privacy of all parties. This method possesses good convergence and scalability, reducing the total operating cost of data center clusters while providing flexible load control capabilities for the power system, supporting the consumption of a high proportion of renewable energy. Attached Figure Description
[0016] Figure 1 This is a simplified flowchart of the distributed collaborative control method for data center clusters in Embodiment 1 of the present invention; Figure 2 This is a schematic diagram of the architecture of the data center cluster distributed collaborative control system in Embodiment 2 of the present invention; Figure 3 This is a schematic diagram of the distributed collaborative optimization process based on the alternating direction multiplier method in Embodiment 2 of the present invention; Figure 4 This is a schematic diagram of the local power-computing power coordination optimization model structure for a single data center in Embodiment 2 of the present invention; Figure 5 This is the convergence graph of the objective function of the distributed iterative solution algorithm in Embodiment 2 of the present invention; Figure 6 This is a line graph showing the 24-hour operating cost distribution of each data center in Embodiment 2 of the present invention; Figure 7 This is a pie chart showing the 24-hour operating cost distribution of each data center in Embodiment 2 of the present invention; Figure 8 This is a time-series variation curve of computing power transmission between data centers in Embodiment 2 of the present invention; Figure 9 This is a line graph showing the local computing power load distribution in each data center in Embodiment 2 of the present invention; Figure 10 This is a comparison chart of the tracking effect between the total computing power requirement and the actual total computing power in Embodiment 2 of the present invention. Detailed Implementation
[0017] The following description, with reference to schematic diagrams, illustrates a distributed collaborative control method and control system for data center clusters according to the present invention. Preferred embodiments of the invention are shown. It should be understood that those skilled in the art can modify the invention described herein while still achieving its advantageous effects. Therefore, the following description should be understood as being of general knowledge to those skilled in the art and is not intended to limit the invention.
[0018] The invention is described more specifically by way of example in the following paragraphs with reference to the accompanying drawings. The advantages and features of the invention will become clearer from the following description. It should be noted that the drawings are in a very simplified form and use non-precise proportions, and are only used to facilitate and clarify the illustration of the embodiments of the invention.
[0019] Various modifications and variations are possible without departing from the spirit and scope of the invention. Thus, if these modifications and variations of the invention fall within the scope of the claims of the invention and their equivalents, the invention also intends to include these modifications and variations.
[0020] Example 1 Please refer to Figure 1 This embodiment discloses a distributed collaborative control method for data center clusters, the method comprising: S1. Establish a collaborative optimization model for data center clusters with the goal of minimizing total system cost; S2. Based on the aforementioned collaborative optimization model, construct local power-computing power coordination optimization sub-models for each data center; S3. Based on the aforementioned sub-model, construct a collaborative optimization model for inter-data center computing load balancing and transmission balancing; S4. Based on the aforementioned collaborative optimization model, the global optimization problem is decomposed into local optimization subproblems that can be solved in parallel and global coordination subproblems; S5. Based on the local optimization subproblem and the global coordination subproblem, design a distributed iterative solution algorithm using the alternating direction multiplier method; S6. Determine whether to continue iterating based on the results of the distributed iterative solution algorithm.
[0021] In this embodiment, a collaborative optimization model integrating electricity costs and computing power transmission costs is established, realizing a two-way collaborative mechanism of "controlling computing with electricity and adjusting electricity with computing power," enabling joint optimal allocation of electricity and computing resources. Employing a distributed solution framework based on the alternating direction multiplier method, each data center only needs to exchange a small amount of boundary information such as local load decisions and computing power transmission plans with the coordinator, without disclosing sensitive data such as internal energy consumption models and customer loads, effectively protecting the commercial privacy of all parties. This method exhibits good convergence and scalability, reducing the total operating cost of data center clusters while providing flexible load regulation capabilities for the power system, supporting the consumption of a high proportion of renewable energy.
[0022] In step S1, the collaborative optimization model includes a minimization objective function and a total computing power balancing constraint function. This step is used to construct a global optimization model, providing a benchmark for distributed collaboration. The model includes a total cost minimization objective and a total computing power balancing constraint.
[0023] The objective function to be minimized is: ; In the formula, t is the index of the scheduling period, t=1,2,......,T; T is the total number of scheduling periods; k is the index of the data center, k=1,2,......,N; N is the number of data centers in the data center cluster; This represents the time-of-use electricity price at time t, expressed in yuan per kilowatt-hour. This represents the electrical power purchased by the k-th data center from the power grid at time t, in kilowatts. This represents the on-grid electricity price at time t, expressed in yuan per kilowatt-hour. This represents the electrical power sold to the grid by the k-th data center at time t, in kilowatts. This represents the transmission price per unit of computing power, expressed in yuan per tera-point operation. This represents the computing power transmitted from the k-th data center to the j-th data center at time t, measured in teraflops per hour.
[0024] The total computing power balance constraint function is: ; In the formula, This represents the local computing load processed by the k-th data center at time t; This represents the total computing power workload required to be completed at time t.
[0025] In step S2, a local optimization model is established for each data center to achieve distributed collaborative optimization while protecting privacy. The power-computing power coordinated optimization sub-model is as follows: ; In the formula, This represents the computing power transmitted from the k-th data center to the j-th data center at time t. The Lagrange multiplier representing the consistency constraint of computing power transmission at time t; This represents the transmission volume published by data center j in the previous iteration; This represents the penalty parameter of the ADMM algorithm. It is the Lagrange multiplier corresponding to the total computing power load balance constraint function, and is the system marginal cost that satisfies the unit total computing power requirement at time t. This indicates the local computing load published by data center j in the previous iteration.
[0026] The constraints of the objective function are: Server utilization constraints: ; Constraints on the range of computing power transmission: ; Computing capacity constraints: ; Power balance constraints: ; In the formula, This represents the server utilization rate at time t; and These represent the minimum and maximum values of server utilization, respectively. This indicates the computing power capacity of the data center; This represents the maximum scaling factor for computing power transfer. This represents the local computing load at time t; The power efficiency coefficient representing the transmission of computing power; This represents the power consumption of the cooling system at time t; This represents the power consumption of the IT equipment at time t.
[0027] In a specific example, each data center runs a local optimizer independently. First, local decision variables and received external variables are initialized. Then, a local optimization problem is constructed, incorporating an economic objective and an ADMM penalty term. The power consumption of IT equipment is modeled using a linear model. The calculation of cooling power consumption uses Model, server utilization Updated in real time.
[0028] In the formula, This represents the base power consumption of a single server in idle state. This indicates the maximum power consumption of a single server when it is fully loaded; Indicates the overall heat transfer coefficient of a data center building; This indicates the total area of the data center server room; Indicates the coefficient of performance of the refrigeration system; This represents the total computing power received from all other data centers at time t.
[0029] Each data center uses an optimization solver to solve the problem in parallel, updating local computing power allocation, power trading, and computing power transmission decisions, providing a foundation for subsequent global coordination.
[0030] In step S3, the collaborative optimization model for computing power load balancing and transmission balancing includes a computing power transmission balancing constraint function and a total computing power load decomposition constraint function: The computing power transmission balance constraint function is as follows: ; ; The total computing power load decomposition constraint is: ; ; In the formula, This represents the total computing load of the k-th data center during time period t. This represents the computing power processed by the k-th data center during time period t.
[0031] In this embodiment, the global constraints are relaxed using the augmented Lagrangian method and embedded as a penalty term into the local objective function of each data center. For the Kth data center, its objective function is reconstructed as: [local economic objective] + [consistency penalty term].
[0032] Among them, the penalty item is ,as well as, This approach transforms the globally coupled constraints, which are difficult to handle directly, into guiding costs that need to be followed in the local optimization of each center, thus laying the foundation for distributed solutions.
[0033] In step S4, the complex coupled optimization model established in S3 is decomposed into a set of simple subproblems that can be solved in parallel. This decomposition is based on the framework of the alternating direction multiplier method, which, by fixing different types of variables, splits the global problem into two alternately solvable parts: local optimization and global coordination.
[0034] Specifically, the expression for the local optimization subproblem is: ; In the formula, Represents the Lagrange multiplier of the computing power transmission consistency constraint at time t during the nth update iteration; This represents the transmission volume published by data center j during the nth update iteration.
[0035] In each iteration, the above embodiments receive the following three types of key information from the central coordinator or obtained through peer-to-peer communication: firstly, the decision results of other data centers regarding computing power transmission. This represents the computing power that other data centers planned to transfer to this center in the previous iteration; secondly, it represents the local load decisions of other data centers. This reflects the amount of local computational tasks undertaken by each center in the previous iteration; thirdly, it reflects the current global coordination signal, namely the Lagrange multipliers. and They respectively encode the current market "price" or "penalty" intensity of the computing power transmission consistency constraint and the total computing power load balancing constraint.
[0036] During each iteration, each data center receives three types of coordination information from the central coordinator or obtains it through point-to-point communication. The first type of information is the computing power transmission decision variables of other data centers. This variable represents the planned amount of computing power transferred from other data centers to this data center in the nth iteration. The second type of information consists of the local load decision variables of other data centers. This variable reflects the amount of local computing tasks allocated to each data center in the nth iteration. The third type of information is the global coordination signal, including the Lagrange multipliers corresponding to the computing power transmission consistency constraints. and the Lagrange multipliers corresponding to the total computing power load balancing constraint The Lagrange multipliers represent the penalty strength of the computing power transmission consistency constraint and the total computing power load balance constraint, respectively.
[0037] Based on the three types of coordination information, the data center independently solves a local optimization problem embedding an augmented Lagrange penalty term. The decision variables for this local optimization problem include: local computing load. Computing power transmitted to each of the other data centers Power purchased from the power grid and the power sold to the grid .
[0038] The constraints of the local optimization problem include: computing resource constraints, which ensure that the sum of local load and received external computing power does not exceed the total computing capacity of the data center, and that the computing power transmission volume is within the allowable range; IT equipment power consumption constraints, which establish a mapping relationship between server utilization and IT equipment power consumption based on a linear model; cooling power consumption constraints, which calculate the power consumption of the cooling system based on IT equipment power consumption and environmental parameters; and power balance constraints, which ensure that instantaneous power balance is met among purchased power, sold power, IT equipment power consumption, cooling system power consumption, and computing power transmission power consumption.
[0039] This optimization problem can typically be modeled as a convex optimization problem and solved using a mathematical programming solver. After the solution is completed, each data center obtains the updated values of the decision variables for the (n+1)th iteration. These updated values include... , , , .in, This represents the local computing load processed by data center j at time t after the (n+1)th iteration update; This represents the computing power transferred from data center k to data center j after the (n+1)th iteration update; This represents the electrical power purchased by data center k from the power grid at time t after the (n+1)th iteration update; This represents the electrical power sold by data center k to the power grid at time t after the (n+1)th iteration update.
[0040] After solving the local optimization subproblems in each data center, the system enters the global coordination and Lagrange multiplier update phase. The technical objective of this phase is to assess the degree of satisfaction of global consistency constraints based on the new decisions generated by the local optimization subproblems, and adaptively adjust the coordination signals accordingly. This global coordination and Lagrange multiplier update phase is executed by a central coordinator, or, in a decentralized architecture, collaboratively across data centers via a predefined distributed communication protocol. The central coordinator first collects the updated decision variables from all data centers in the (n+1)th iteration. These decision variables include the local computing load of each data center. and computing power transmission decision Subsequently, the central coordinator synchronously updates the Lagrange multipliers according to the update rules of the alternating direction multiplier method.
[0041] The global coordination subproblems include the computing power transmission consistency multiplier update problem and the total computing power load balancing multiplier update problem.
[0042] The calculation method for the consistency multiplier update problem of computing power transmission is as follows: ; In the formula, This represents the Lagrange multiplier of the consistency constraint for computing power transmission at time t during the (n+1)th update iteration; The calculation method for the total computing power load balancing multiplier update problem is as follows: ; In the formula, and These represent the Lagrange multipliers corresponding to the total computing power load balance constraint function at the (n+1)th and nth update iterations, respectively; This represents the local computing load processed by the k-th data center at time t during the (n+1)-th update iteration.
[0043] In this embodiment, when there is a deviation between the total supply and total demand of computing power in the system, the Lagrange multiplier corresponding to the system balance will be adjusted according to the supply and demand difference, thereby incentivizing all data centers to coordinately adjust their local computing power load in subsequent iterations, driving the total supply of the system to tend towards the balance of total demand.
[0044] In step S5, based on the local optimization subproblem and the global coordination subproblem, the distributed iterative solution algorithm of the alternating direction multiplier method is designed, including: S51. Based on the Lagrange multipliers of the transmission consistency constraint and the Lagrange multipliers corresponding to the total computing power load balancing constraint function calculated based on the global coordination subproblem, and the decision variables calculated in the local optimization subproblem, the coordination variables and decision variables are initialized; S52. Each data center solves its own local optimization problem based on the coordination variable and the decision variable, and updates its local decision variable accordingly; S53. Update the coordination variable based on the local decision variable, and send the updated coordination variable to each of the data centers.
[0045] In this embodiment, the coordinating variable is the Lagrange multiplier. Lagrange multipliers corresponding to the total computing power load balancing constraint .
[0046] The purpose of this step is to transform the established optimization model and decomposition framework into a concrete, executable, and convergent distributed collaborative solution algorithm. This algorithm, through a finite number of iterations and limited information interaction among multiple agents, gradually coordinates their respective decisions, ultimately approximating the optimal operating state of the system.
[0047] Specifically, in step S51, the initialization of the coordination and decision variables forms the basis for the iterative solution. First, the maximum number of iterations is defined. As one of the stopping conditions of the algorithm, it prevents infinite loops in the case of non-convergence. Subsequently, all global coordination variables are initialized: the Lagrange multipliers corresponding to the computing power transfer consistency constraint are set. Lagrange multipliers corresponding to the total computing power load balancing constraint Setting it to zero indicates that the system has not yet imposed any "penalty" for constraint violations in the initial state; this applies to the decision-making regarding the transfer of computing power between all data centers. Local computing load decisions for each data center Set it to zero or assign a reasonable initial value based on historical data. Simultaneously, the penalty parameter ρ for the ADMM algorithm needs to be set to >0. This parameter balances the importance of the original objective function and the constraint violation penalty term; its value affects the convergence speed and usually needs to be determined through experimental debugging.
[0048] In step S52, after initialization, the algorithm enters an iterative loop. Each data center solves its own local optimization subproblem based on the coordination variable and the decision variable, and updates its local decision variable. In each iteration, all data centers simultaneously and independently solve their own local optimization subproblem based on the global coordination variable generated in the previous iteration or initialization, obtaining the updated local decision variable.
[0049] In step S53, following the local optimization phase, global information exchange is required to achieve system collaboration. Each data center sends the key decision variables calculated in this round to the central coordinator. If the system adopts a fully distributed architecture, each center sends relevant information to other data centers with which it has interaction via point-to-point communication.
[0050] In this embodiment, step S6 includes: After receiving the updated local decision variables from each of the data centers, it is determined whether the computing power transmission consistency error and the total computing power load balancing error of each of the data centers are both less than the preset convergence threshold and whether the number of iterations has reached the maximum number of iterations. If the convergence condition is met and the maximum number of iterations is reached, the iteration is terminated; otherwise, the local optimization subproblem and the global coordination subproblem are solved, and the consistency error of computing power transmission and the total computing power load balancing error of each data center are re-evaluated to determine whether the iteration termination condition is met.
[0051] The central coordinator calculates the new round of Lagrange multipliers according to the update rules of the alternating direction multiplier method. These Lagrange multipliers include those corresponding to the computing power transmission consistency constraint. and The updated coordination variables are broadcast to all data centers for solving the local optimization subproblems in the next iteration.
[0052] Example 2 This embodiment uses a cluster consisting of three heterogeneous data centers for simulation verification. The system architecture is as follows: Figure 2 As shown, the system comprises three data centers (DC1, DC2, and DC3), a regional power grid, a coordinator, and a central dispatch center for total computing power demand. The three data centers are interconnected via a computing power transmission network, each trading electricity with the regional power grid, and achieving distributed collaborative optimization through the coordinator. DC1 is equipped with a photovoltaic power generation system, DC2 has the largest computing capacity, and DC3 is a compact data center. The three data centers differ in capacity, energy efficiency, and cost characteristics, simulating the heterogeneous characteristics of a real data center cluster.
[0053] Using the time-of-use electricity prices shown in Table 1 and Table 2, the three data centers differ in computing capacity, power consumption characteristics, and energy efficiency parameters to simulate a real-world heterogeneous data center cluster. The simulation environment is based on the MATLAB R2024b platform, using the YALMIP toolbox for optimization modeling and the Gurobi solver for solving the problem. The hardware configuration is an Intel Core i7-11800H processor and 32GB of RAM. The scheduling cycle is set to 24 hours, and the time resolution is 1 hour. In the ADMM algorithm parameters, the penalty parameter ρ is set to 1.0, and the maximum number of iterations is set to 50.
[0054] Table 1 Table 2 The complete process of the ADMM distributed collaborative optimization algorithm is as follows: Figure 3 As shown. After the algorithm starts, initialization is performed first, setting the iteration counter k=0, and all Lagrange multipliers... , Initialized to 0, the computing power transfer variable between data centers Initialize to 0. After initialization, enter the parallel local optimization phase, where each data center solves independently, as follows: Figure 3 The local optimization problem shown.
[0055] like Figure 4 As shown, the local optimization model for a single data center comprises four parts: an input layer, an optimization model layer, a solution layer, and an output layer. Input parameters include external parameters (electricity price, computing power demand, etc.), collaborative parameters (Lagrange multipliers, decisions made by other data centers), and local parameters (capacity, efficiency, photovoltaics, etc.). The objective function of the optimization model layer is min(electricity cost + computing power transmission cost + ADMM penalty term), and the constraints include computing power capacity constraints, computing power transmission range constraints, server utilization constraints, IT equipment power consumption models, cooling power consumption models, and power balance constraints.
[0056] like Figure 5 The convergence line graph shown demonstrates that the ADMM algorithm exhibits good convergence performance. The total computational tracking error decreased from the initial 23.57 TFLOPS to 0 TFLOPS.
[0057] The results of the economic analysis are as follows Figure 6 and Figure 7 As shown, the cost distribution of the three data centers over 24 hours exhibits distinct spatiotemporal characteristics: all three data centers have lower costs, even negative costs, during the daytime (0:00-8:00); DC2 has the highest cost, which increases during peak hours (8:00-9:00, 16:00-17:00); DC1's cost falls between the two; and DC3 has the lowest cost.
[0058] The effect of computing power scheduling is as follows Figure 8 , Figure 9 and Figure 10 As shown. Figure 8 It shows the temporal changes in computing power transmission between data centers, with six curves corresponding to L12, L13, L21, L23, L31, and L32, respectively.
[0059] Figure 10 The diagram illustrates the total computing power allocation and tracking performance. The left graph shows the total computing power requirement and the actual total computing power, while the right graph, a stacked area diagram, shows the timing allocation of L1, L2, and L3. It is evident that the actual total computing power closely tracks the demand curve.
[0060] In summary, the data center cluster distributed collaborative control method is based on a "global-local" two-layer optimization framework built upon the data center cluster, forming a multi-data center collaborative scheduling system. At the global level, the total computing power load balance constraint is established with the goal of minimizing the total system cost. At the local level, the data centers are modeled, integrating electricity costs and computing power transmission costs, and coupling power balance with computing power load constraints to achieve a two-way collaborative mechanism of "controlling computing with electricity and adjusting electricity with computing power." Computing power transmission decisions directly affect power balance, and real-time electricity price signals adjust computing power scheduling inversely, achieving optimal resource allocation in the spatiotemporal dimension. The collaborative process adopts an alternating direction multiplier method distributed solution framework, with each data center solving its local optimization problem in parallel. It only needs to exchange a small amount of boundary information, such as planned computing power and local load decisions, with the coordinator, without disclosing sensitive data such as internal energy consumption models and customer load details. This invention is applicable to large-scale cluster collaborative operation scenarios that include heterogeneous data centers. It can provide power systems with flexible and reliable load regulation capabilities, support the consumption of a high proportion of renewable energy, and provide key technical support for data center operators to reduce operating costs and improve service reliability. It has important theoretical value and engineering application prospects.
[0061] Example 3 Based on the same inventive concept, this embodiment discloses a distributed collaborative control system for data center clusters. The system is used to implement the distributed collaborative control method for data center clusters disclosed in Embodiments 1 and 2. The system includes: The collaborative optimization model building module is used to build a collaborative optimization model for data center clusters with the goal of minimizing the total system cost.
[0062] The local sub-model construction module is used to construct local power-computing power coordination optimization sub-models for each data center based on the collaborative optimization model.
[0063] The collaborative optimization model construction module is used to construct a collaborative optimization model for computing load balancing and transmission balancing between data centers based on the sub-model.
[0064] The problem decomposition module is used to decompose the global optimization problem into local optimization subproblems and global coordination subproblems that can be solved in parallel, based on the collaborative optimization model.
[0065] The distributed iterative solution module is used to design a distributed iterative solution algorithm using the alternating direction multiplier method based on the local optimization subproblem and the global coordination subproblem.
[0066] The iteration judgment module is used to determine whether to continue iterating based on the result of the distributed iterative solution algorithm.
[0067] Specifically, the collaborative optimization model establishment module is used to establish a minimum objective function; comprehensively consider the electricity purchase cost, electricity sales revenue and computing power transmission cost between data centers; and establish a total computing power balance constraint function so that the sum of the local computing power load processed by all data centers is equal to the total computing power task requirement that the system needs to complete at time t.
[0068] The local sub-model construction module establishes an independent local optimization model for each data center based on the collaborative optimization model, achieving distributed collaborative optimization while protecting privacy. This module embeds the computing power transmission consistency constraint and the total computing power load balancing constraint into the objective function using the augmented Lagrangian method, forming a local optimization problem that includes an economic objective and an ADMM penalty term. It also sets server utilization constraints, computing power capacity constraints, computing power transmission range constraints, power balance constraints, and a cooling power consumption model to establish a mapping relationship between computing power load and power consumption.
[0069] The collaborative optimization model construction module establishes a computing power transmission balance constraint function to achieve bidirectional consistency of computing power transmission between data centers, i.e., the amount of transmission and reception are equal; it establishes a total computing power load decomposition constraint function to decompose the total computing power load of each data center into the sum of local processing computing power and the transmission computing power received from other data centers; and it relaxes the global coupling constraints through the augmented Lagrange method and embeds them as a penalty term into the local objective function of each data center to realize a bidirectional collaborative mechanism of "controlling computing with electricity and adjusting electricity with computing".
[0070] The problem decomposition module is based on the alternating direction multiplier method framework. By fixing different types of variables, it breaks down the global problem into two alternately solvable parts: local optimization and global coordination. The local optimization subproblem (subproblem P1) is defined as an optimization problem that each data center solves independently under the premise of fixing the decision variables of other data centers and the global Lagrange multipliers. The decision variables include local computing power load, computing power transferred to other data centers, purchased electricity power, and sold electricity power. The global coordination subproblem (subproblem P2) includes computing power transmission consistency multiplier update and total computing power load balancing multiplier update.
[0071] The distributed iterative solution module first performs algorithm initialization, defining the maximum number of iterations, setting all Lagrange multipliers to zero, setting the computing power transmission decision and local computing power load decision to initial values, and setting the penalty parameters for the ADMM algorithm. Then, it enters the parallel local optimization stage, where each data center solves its own local optimization problem simultaneously and independently based on the global variables generated in the previous iteration. It uses the YALMIP toolbox for optimization modeling and calls the Gurobi solver to solve the problem, updating the local decision variables. Finally, it performs global information exchange and coordination, where each data center sends the key decision variables calculated in this round to the central coordinator or to other data centers with interactive relationships via point-to-point communication. The coordinator calculates the new round of Lagrange multipliers according to the ADMM update rules and broadcasts them to all data centers.
[0072] After receiving the updated local decision variables from each data center, the iterative judgment module calculates the computing power transmission consistency error and the total computing power load balancing error, and determines whether the above errors are all less than the preset convergence threshold and whether the number of iterations has reached the maximum number of iterations. If the convergence condition is met or the maximum number of iterations is reached, the iteration is terminated and the optimal scheduling scheme is output, including the local computing power allocation of each data center, the power trading strategy, and the computing power transmission plan between data centers. Otherwise, it returns to the distributed iterative solution module to continue to execute the next round of iterations until the algorithm converges.
Claims
1. A distributed collaborative control method for data center clusters, characterized in that, The method includes: Establish a collaborative optimization model for data center clusters with the goal of minimizing total system cost; Based on the aforementioned collaborative optimization model, a local power-computing power coordination optimization sub-model is constructed for each data center. Based on the aforementioned sub-model, a collaborative optimization model for inter-data center computing load balancing and transmission balancing is constructed. Based on the aforementioned collaborative optimization model, the global optimization problem is decomposed into local optimization subproblems that can be solved in parallel and global coordination subproblems. Based on the local optimization subproblem and the global coordination subproblem, a distributed iterative solution algorithm using the alternating direction multiplier method is designed. The decision to continue iteration is based on the results of the distributed iterative solution algorithm.
2. The distributed collaborative control method for data center clusters as described in claim 1, characterized in that, The collaborative optimization model includes a minimization objective function and a total computing power balance constraint function: The objective function to be minimized is: ; In the formula, t is the index of the scheduling period, t=1,2,......,T; T is the total number of scheduling periods; k is the index of the data center, k=1,2,......,N; N is the number of data centers in the data center cluster; This represents the time-of-use electricity price at time t, expressed in yuan per kilowatt-hour. This represents the electrical power purchased by the k-th data center from the power grid at time t, in kilowatts. This represents the on-grid electricity price at time t, expressed in yuan per kilowatt-hour. This represents the electrical power sold to the grid by the k-th data center at time t, in kilowatts. This represents the transmission price per unit of computing power, expressed in yuan per tera-point operation. This represents the computing power transmitted from the k-th data center to the j-th data center at time t, measured in teraflops per hour. The total computing power balance constraint function is: ; In the formula, This represents the local computing load processed by the k-th data center at time t; This represents the total computing power required to complete tasks at time t.
3. The distributed collaborative control method for data center clusters as described in claim 2, characterized in that, The power-computing power coordinated optimization sub-model is as follows: ; In the formula, This represents the computing power transmitted from the k-th data center to the j-th data center at time t. The Lagrange multiplier representing the consistency constraint of computing power transmission at time t; This represents the transmission volume published by data center j in the previous iteration; This represents the penalty parameter of the ADMM algorithm. It is the Lagrange multiplier corresponding to the total computing power load balance constraint function, and is the system marginal cost that satisfies the unit total computing power requirement at time t; This indicates the local computing load published by data center j in the previous iteration.
4. The distributed collaborative control method for data center clusters as described in claim 3, characterized in that, The constraints of the objective function are: Server utilization constraints: ; Constraints on the range of computing power transmission: ; Computing capacity constraints: ; Power balance constraints: ; In the formula, This represents the server utilization rate at time t; and These represent the minimum and maximum values of server utilization, respectively. This indicates the computing power capacity of the data center; This represents the maximum scaling factor for computing power transfer. This represents the local computing load at time t; The power efficiency coefficient representing the transmission of computing power; This represents the power consumption of the cooling system at time t; This represents the power consumption of the IT equipment at time t; in, ; ; ; In the formula, This represents the base power consumption of a single server in idle state. This indicates the maximum power consumption of a single server when it is fully loaded; Indicates the overall heat transfer coefficient of a data center building; This indicates the total area of the data center server room; Indicates the coefficient of performance of the refrigeration system; This represents the total computing power received from all other data centers at time t.
5. The distributed collaborative control method for data center clusters as described in claim 1, characterized in that, The collaborative optimization model for computing power load balancing and transmission balancing includes a computing power transmission balancing constraint function and a total computing power load decomposition constraint function: The computing power transmission balance constraint function is: ; in, ; The total computing power load decomposition constraint is: ; in, ; In the formula, This represents the total computing load of the k-th data center during time period t. This represents the computing power processed by the k-th data center during time period t.
6. The distributed collaborative control method for data center clusters as described in claim 1, characterized in that, The expression for the local optimization subproblem is: ; In the formula, Represents the Lagrange multiplier of the consistency constraint of computing power transmission at time t during the nth update iteration; This represents the transmission volume published by data center j during the nth update iteration.
7. The distributed collaborative control method for data center clusters as described in claim 1, characterized in that, The global coordination subproblems include the computing power transmission consistency multiplier update problem and the total computing power load balancing multiplier update problem; The calculation method for the computing power transmission consistency multiplier update problem is as follows: ; In the formula, This represents the Lagrange multiplier of the consistency constraint for computing power transmission at time t during the (n+1)th update iteration; The calculation method for the total computing power load balancing multiplier update problem is as follows: ; In the formula, and These represent the Lagrange multipliers corresponding to the total computing power load balance constraint function at the (n+1)th and nth update iterations, respectively; This represents the local computing load processed by the k-th data center at time t during the (n+1)-th update iteration.
8. The distributed collaborative control method for data center clusters as described in claim 7, characterized in that, Based on the local optimization subproblem and the global coordination subproblem, a distributed iterative solution algorithm using the alternating direction multiplier method is designed, including: The Lagrange multipliers corresponding to the transmission consistency constraint and the total computing power load balancing constraint function calculated based on the global coordination subproblem, as well as the decision variables calculated in the local optimization subproblem, are used to initialize the coordination variables and decision variables. Each data center solves its own local optimization problem based on the coordination variable and the decision variable, and updates its local decision variable accordingly. The coordination variable is updated based on the local decision variable, and the updated coordination variable is sent to each of the data centers.
9. The distributed collaborative control method for data center clusters as described in claim 1, characterized in that, Determining whether to continue iteration based on the results of the distributed iterative solution algorithm includes: After receiving the updated local decision variables from each of the data centers, determine whether the computing power transmission consistency error and the total computing power load balancing error of each of the data centers are both less than the preset convergence threshold or whether the number of iterations has reached the maximum number of iterations. If the convergence condition is met or the number of iterations reaches the maximum number of iterations, the iteration is terminated; otherwise, the local optimization subproblem and the global coordination subproblem are solved, and the consistency error of computing power transmission and the total computing power load balancing error of each data center are re-evaluated to determine whether the iteration termination condition is met.
10. A distributed collaborative control system for a data center cluster, characterized in that, The system includes: The collaborative optimization model building module is used to build a collaborative optimization model for data center clusters with the goal of minimizing the total system cost. The local sub-model construction module is used to construct local power-computing power coordination optimization sub-models for each data center based on the collaborative optimization model. The collaborative optimization model construction module is used to construct a collaborative optimization model for computing load balancing and transmission balancing between data centers based on the sub-model. The problem decomposition module is used to decompose the global optimization problem into local optimization subproblems and global coordination subproblems that can be solved in parallel, based on the collaborative optimization model. The distributed iterative solution module is used to design a distributed iterative solution algorithm using the alternating direction multiplier method based on the local optimization subproblem and the global coordination subproblem. The iteration judgment module is used to determine whether to continue iterating based on the result of the distributed iterative solution algorithm.