Data Center Control Hierarchy for Neural Network Ensemble
By introducing a control architecture of the load part and resource part in the data center system, using neural network models to predict the required power is solved, and the challenges of data center systems in handling different workloads and improving energy efficiency are achieved, achieving efficient power management and system scalability.
Patent Information
- Application Number
- CN202111353871.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-03-18
- Filing Date
- 2021-11-10
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2041-11-10
AI Technical Summary
Existing data center systems have challenges in handling different types of workloads and improving energy efficiency, and conventional control architectures lack scalability and technical reusability.
The control architecture is adopted that includes a load part and a resource part. The load part is composed of an electronic rack, a thermal management system and a power flow optimizer. The power flow optimizer uses a neural network model to predict the amount of power required based on thermal data and load data; the resource part is configured and selected through the resource controller to provide power.
It realizes efficient power management of data center systems under different workload conditions, improving energy efficiency and system scalability.
Smart Images

Figure CN115113705B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention generally relate to data center systems. More specifically, embodiments of the present invention relate to a control architecture for a data center system. Background Art
[0002] With the rapid development of AI (Artificial Intelligence), big data, edge computing, etc., the demand for data centers and IT (Information Technology) clusters has become increasingly challenging. The challenges faced not only lie in the sharp increase in the number of data centers and servers to be deployed, but also in the differences between different types of workload demands. These demands are the main driving factors for the rapid development of data centers. However, this challenge requires data centers to be able to adapt to workload changes. Since workload changes are directly related to the diversity of IT servers, more challenging is that energy efficiency is always one of the requirements for data centers and IT clusters. Energy efficiency is not only related to power consumption and Opex, but more importantly, it meets environmental and power usage regulations.
[0003] Another challenge is that the control design of data centers is complex. Since there are completely different control technical fields based on different systems, and they are also closely connected to each other during normal operation. It is important to combine them organically.
[0004] AI and ML (Machine Learning) technologies will sooner or later become key tools and technologies for data centers and IT clusters. It will have a comprehensive impact on data centers, including design, construction, deployment, and operation. It can bring various benefits to the intelligent control of data centers. The current challenge is that data centers generate a large amount of data. Completing the model training and adjustment of the cluster is expensive and time-consuming. Given the nature of the data center system, a well-trained model based on one cluster may be well-suited for that cluster. However, it may perform poorly in another cluster, or may require a large amount of retraining. It may be applicable to another cluster that is the same but connected to different systems (e.g., cooling and power).
[0005] Conventional solutions for designing data center control include separate modules, such as a control module for the cooling system, a control module for the power system, a control module for IT, and various modules that may be used for IT control. All these control modules may not be fully integrated to achieve a joint design. The disadvantage is that it is extremely complex to integrate them organically and operate them as a complete system. Generally speaking, conventional solutions lack scalability and technical reusability. Summary of the Invention
[0006] One aspect of the present disclosure relates to a data center system, including: a load part having a plurality of electronic racks, a thermal management system, and a power flow optimizer, wherein each of the electronic racks includes a plurality of servers, and each server contains one or more electronic devices, wherein the thermal management system is configured to provide cooling and / or heating to the electronic devices, and wherein the power flow optimizer is configured to determine the load power demand of the load part based on the thermal data of the thermal management system and the load data of the electronic racks; and a resource part having a plurality of power supplies that supply power to the load part, wherein the resource part includes a resource controller that configures and selects at least some of the power supplies to supply power to the load part based on the load power demand provided by the power flow optimizer, wherein the power flow optimizer includes a power flow neural network model for predicting the amount of power required by the electronic racks and the thermal management system based on the thermal data and the load data to meet the thermal demand and the data processing load demand of the load part.
[0007] Another aspect of the present disclosure relates to a method for managing a data center system, the method including: using a power flow optimizer to determine the load power demand of a load part having a thermal management system and a plurality of electronic racks based on the thermal data of the thermal management system and the load data of the electronic racks, wherein each of the electronic racks includes a plurality of servers, and each server contains one or more electronic devices, wherein the thermal management system is configured to provide cooling and / or heating to the electronic devices; and configuring and selecting at least some of the power supplies by a resource controller of a resource part having a plurality of power supplies to supply power to the load part based on the load power demand provided by the power flow optimizer, wherein the power flow optimizer includes a power flow neural network model for predicting the amount of power required by the electronic racks and the thermal management system based on the thermal data and the load data to meet the thermal demand and the data processing load demand of the load part. Brief Description of the Drawings
[0008] Embodiments of the present invention are shown by way of example and not limitation in the figures of the accompanying drawings, wherein like reference numerals represent like elements.
[0009] Figure 1 is a block diagram showing the overall architecture of a data center system according to one embodiment.
[0010] Figure 2 is a flowchart showing the process of the load part of a data center system according to one embodiment.
[0011] Figure 3 is a flowchart showing the process of managing power supplies according to one embodiment.
[0012] Figure 4 is a flowchart showing the process of operating an intermediate level according to one embodiment.
[0013] Figure 5 shows an overall three - level control hierarchy design and operation method according to one embodiment.
[0014] Figure 6A and Figure 6B shows a larger - scale system with a plurality of subsystems interconnected with each other according to one embodiment.
[0015] Figure 7 is a block diagram showing a multi - cluster design of a power system according to one embodiment.
[0016] Figure 8 is a flowchart showing a process of managing the power of a data center according to one embodiment. Detailed Embodiments
[0017] Various embodiments and aspects of the present invention will be described with reference to the details discussed below, and the accompanying drawings will illustrate the various embodiments. The following description and the drawings are illustrative of the present invention and should not be construed as limiting the present invention. Many specific details are described to provide a thorough understanding of the various embodiments of the present invention. However, in some cases, well - known or conventional details are not described in order to provide a concise discussion of the embodiments of the present invention.
[0018] Reference to "one embodiment" or "an embodiment" in the specification means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of the present invention. The phrase "in one embodiment" appearing in various places in the specification does not necessarily refer to the same embodiment.
[0019] Embodiments of the present disclosure relate to a control hierarchy design for a data center. First, the system design and control flow are introduced. A high - level description of the main components in the system is provided in this part, including electrical components, mechanical components, and IT components and their interconnections in the entire system. Then, the control flow is introduced to present the overall system management. There are three levels in the system, namely the load level, the resource level, and the intermediate level. In the second part, the design and control flow of each level are introduced to provide a detailed view of each level. Input / output is given to show the control logic in each level. This is intended to provide a decoupled design in terms of control while still keeping the entire system as an organic system joined together. An optimizer is used in combination with each controller to assist in implementing AI / ML models. The controller location and functions, as well as the operation design of different levels, are introduced in detail.
[0020] According to some embodiments, a data center system includes a load portion having an array of electronic racks, a thermal management system, and a power flow optimizer. Each of the electronic racks includes a stack of servers, and each server contains one or more electronic devices. The thermal management system is configured to provide cooling and / or heating to the electronic devices. The power flow optimizer is configured to determine the load power demand of the load portion based on the thermal data of the thermal management system and the load data of the electronic racks. The data center system further includes a resource portion having a plurality of power supplies that supply power to the load portion. The resource portion includes a resource controller that configures and selects at least some of the power supplies based on the load power demand provided by the power flow optimizer to supply power to the load portion. The power flow optimizer includes a power flow neural network (NN) model for predicting the amount of power required by the electronic racks and the thermal management system based on the thermal data and the load data to meet the thermal demand and the data processing load demand of the load portion.
[0021] In one embodiment, the data center system further includes an intermediate portion coupled between the resource portion and the load portion, wherein the intermediate portion includes a power bus to distribute power from the resource portion to the load portion and other subsystems. The intermediate portion further includes: a subsystem load detector coupled to the other subsystems to determine the subsystem power demand; and a central controller coupled to the subsystem load detector and the power flow optimizer of the load portion to determine the total power demand based on the subsystem power demand and the load power demand. The resource controller utilizes the total power demand to configure and select at least some of the power supplies.
[0022] In one embodiment, the load portion further includes: one or more temperature sensors disposed within each server to measure the temperature of the electronic devices; and a workload detector configured to determine the workload of each of the servers. The power flow NN model infers the load power demand based on the temperature and the workload of each of the servers. The load portion further includes a power scheduling controller coupled to the power flow controller to proportionally distribute the power received from the resource portion to the thermal management system and the servers based on the load power demand received from the power flow controller.
[0023] In one embodiment, the load power demand includes information on how to schedule power to the thermal management system and the servers. The power scheduling controller is configured to output the total power required by the load portion to the central controller. The resource portion includes a resource optimizer for receiving the total power demand from the central controller to generate power supply configuration information. The resource controller is configured to configure the power supplies based on the power supply configuration information. In one embodiment, power optimization on the load side including the NN model can be developed to achieve optimized computational efficiency.
[0024] In one embodiment, the resource optimizer includes an NN model to determine power supply configuration information based on the total power demand. The power supply configuration information includes information specifying the amount of power to be provided by each of a plurality of power supplies. The power supplies include a utility power supply, a photovoltaic (PV) power supply, and a battery power supply. The central controller includes an NN model for inferring the total power demand based on the load power demand and the subsystem power demand. The data center system is the first data center subsystem among a plurality of data center subsystems. The power bus in the middle part is connected to the power bus in the middle part of the second data center subsystem of the data center subsystem.
[0025] In one embodiment, the central controller is shared by the first data center subsystem and the second data center subsystem. The data center subsystem is part of the first data center cluster of the data center cluster. Each in the data center cluster is controlled by a corresponding cluster controller, and among them, the central controller is shared by a plurality of data center clusters.
[0026] Figure 1 is a block diagram showing the overall architecture of a data center system according to one embodiment. Referring to Figure 1 , the data center configuration or architecture 100 can represent any data center, where the data center can include one or more arrays of electronic racks. Each electronic rack includes one or more server chassis arranged in a stacked manner. Each server chassis includes one or more servers operating therein. Each server may include one or more processors, memories, storage devices, network interfaces, etc., collectively referred to as IT components. In addition, the data center may also include a thermal management system to provide cooling to the IT components that generate heat during operation. Data center cooling may include liquid cooling and / or air cooling.
[0027] In one embodiment, the data center architecture 100 includes a resource part or resource level 101, an intermediate part or intermediate level 102, and a load part or load level 103. The load level 103 includes an IT load 112 and a thermal management system (also referred to as thermal management for all levels) 113. The IT load 112 may represent one or more electronic racks, each containing one or more servers. The thermal management system 113 may provide liquid cooling and / or air cooling to the IT components of the servers. In one embodiment, some of the IT components (e.g., processors) may be attached to cold plates for liquid cooling and / or attached to radiators for air cooling. In addition, the load level 103 also includes one or more temperature sensors 114 for measuring the temperature at different positions within the load level (e.g., the surface of the IT components, the temperature of the cooling liquid, the ambient temperature, etc.). The load level 103 also includes a load detector 115 for determining or detecting the workload of the load 112, which may be proportional to the power consumption of the load 112.
[0028] Load level 103 includes IT load 112 and thermal management 113 at all levels from within the server (such as cold plates, TEC (thermoelectric cooling)) to the overall system level. At this level, the critical connection is temperature, which is measured by one or more temperature sensors 114. This means that temperature is used to connect the entire system between IT and cooling. However, another key input at this level is the workload, which can be determined or detected by a load detector 115. Therefore, the workload is also used to design the control at this level. Temperature is considered a dependent factor of the load; however, it is also closely related to the thermal system. In one embodiment, the load detector 115 is connected to switch logic provided on the server and / or an electronic rack (e.g., a motherboard) to determine the workload and traffic passing through the network interface. In one embodiment, the load detector 115 is connected to the respective BMCs (board management controllers) of the server chassis to determine the workload of various components, such as processor usage, etc. In some architectures, there are load balancing servers or resource managers for dispatching the workload to individual servers, and then the load detector 115 can receive information about the distributed workload from these components.
[0029] In addition, load level 103 includes a power flow optimizer 111, which can be implemented as a processor, a microcontroller, an FPGA (field programmable gate array), or an ASIC (application specific integrated circuit). The power flow optimizer 111 is configured to determine the load power demand of load level 103 based on thermal data (e.g., temperature) provided by the temperature sensors 114 and load data provided by the load detector 115. In one embodiment, the power flow optimizer 111 includes a machine learning model such as a neural network (NN) model to predict or determine the load power demand based on temperature data and load data. The load power demand represents the amount of power required by the IT load 112 and the thermal management system 113 to meet the thermal requirements (e.g., operate below a predetermined temperature) and data processing load requirements of the IT load 112. A large amount of thermal data and load data at various time points for various loads can be used to train the NN model. The NN model is configured to infer the load power demand based on the temperature data and workload of the server. In one embodiment, the optimized power demand generated by the power flow optimizer 111 includes the optimized power demand for each individual server (i.e., at the server level) and / or the associated thermal management system.
[0030] In one embodiment, the load level 103 (also referred to as level 0) further includes a scheduling controller 110 for receiving load power demand information from the power flow optimizer 111. The power demand information may include information on how to schedule or allocate power to the IT load 112 and the thermal management system 113, where power is received from the resource level 101 via the intermediate level 102, which will be described in further detail below. The scheduling controller 110 may control or configure the switching logic (e.g., S4, S5), as shown by the dashed lines, to control and allocate appropriate power to the IT load 112 and the thermal management system 113 based on the load power demand information provided by the power flow optimizer 111. The scheduling controller 110 also provides the load power demand information to the central controller 109 of the intermediate level 102.
[0031] In the load level 103, temperature is used as a key parameter for the load and the thermal system. The load detection performed by the load detector 115 plays a significant role. The load detector 115 (or the power flow optimizer 111) receives the actual workload and converts it into the actual power required by the workload. Additionally, in a more advanced architecture, the load detection also provides an optimized workload allocation strategy. Since the load does not directly reflect the thermal system, temperature is used to connect the load and the thermal system. Temperature and load detection are used as inputs to the power flow optimizer 111. The NN model of the power flow optimizer 111 only inputs these two parameters and produces an output representing the load power demand.
[0032] A training dataset can be used to train the NN model. Once the training dataset (such as the temperature range and the load power range) converges well, the power flow optimizer 111 can provide more accurate power scheduling at this level, such as the cooling power to the thermal management system 113 and the load power to the IT load 112. Note that the load power to the IT load 112 is different from the computed power because the load power to the IT load 112 may be greater than the computed power due to power loss and power leakage. The thermal management may affect this difference, thus changing the corresponding required power to the thermal management system 113. All these complex strategies are implemented by the NN model in the power flow optimizer 111. However, the only output of the scheduling controller is the total required load power.
[0033] Figure 2 is a flowchart showing the process of the load part of a data center system according to one embodiment. The process 200 may be executed by Figure 1 the load level 103. Referring to Figure 2, at block 201, the load detector 115 determines the workload of the IT load 112 and may convert the load data into a power demand. Additionally, the temperature sensor 114 measures the temperature associated with the IT load 112. At block 202, the temperature data and the load data are fed into the input of the NN model of the power flow optimizer 111, which results in an optimized power demand schedule for the thermal system and the load. At block 203, in response to the load power demand, the scheduling controller 110 controls the switching logic to supply appropriate power to the thermal system and the load. Note that the term "load power demand" refers to the power demand of the load portion or load level 103, including the power demand of the IT load 112 and the thermal management system 113. At block 204, the scheduling controller 110 outputs a request for the load power to the central controller 109.
[0034] Back to Figure 1 , in one embodiment, the resource portion or resource level 101 includes various power sources or energy sources, such as a utility power source 104, a photovoltaic (PV) power source 105, a storage power source 107 (e.g., a battery), and other energy sources 106. The utility power source 104 supplies alternating current (AC) power from a public power grid (e.g., provided by a public utility company), and this AC power can be converted into direct current (DC) power using an AC to DC (AC / DC) converter. The PV power source 105 can be a DC power source, which can be converted into a different DC power voltage using a DC to DC (DC / DC) converter. The storage power source 107 can be charged by any one of the power sources 104 - 106. When the power supplied by the power sources 104 - 106 is insufficient, the storage power source 107 can discharge to supply power to the load portion 103.
[0035] In one embodiment, the resource level 101 includes a resource controller 116 for configuring and selecting at least some of the power sources 104 - 107 to supply power to the load level 103 based at least on a lower power demand. The resource controller 116 controls the switching logic, as shown by the dashed line, to configure and select the power sources.
[0036] In one embodiment, the resource level 101 further includes a resource optimizer 108 for optimizing and generating power source configuration information. The power source configuration information includes selection information for selecting at least some of the power sources 104 - 107. The resource controller 116 utilizes the power source configuration information to control the power sources 104 - 107. In one embodiment, the resource optimizer 108 includes an NN model to determine the power source configuration information based on the total required power. The power source configuration information can include information indicating the amount of power to be supplied by each of the power sources 104 - 107.
[0037] At resource level 101, it is shown that this level is mainly designed for energy. It can be seen that there are several different types of power sources, including utility power source 104, PV power source 105, and other energy sources 106. In addition, backup energy or storage power source 107 is used at this level. The resource controller 116 is used to control the switching logic (S1, S2, S3) to connect the power to the main source bus. The resource optimizer 108 is used to provide a scheduling strategy and communicate with the central controller 109. In one embodiment, the resource optimizer 108 includes an NN model for optimizing power distribution based on the required total power and the existing power conditions and availability from each power source. The only input fed into the resource optimizer 108 is the required total power. All other variations, namely different power availabilities and conditions, are also inputs, but may not need to be considered variables.
[0038] At this level, the only input from the outside is the required total power. It can be the actual power in kW or kWh or a dimensionless value representing the power demand. The resource optimizer 108 is integrated with an AI / ML model to provide the most effective inference on detailed power scheduling. The scheduling strategy is passed to the resource controller 116, and the resource controller 116 manages the power inputs from the utility, PV system, other renewable power sources, batteries, etc. Therefore, it can be seen that this level is highly decoupled from other levels.
[0039] Since the total power is the only input. This is beneficial for the NN model because the variation in the input is only the total power, which can be easily covered by the training dataset. On the resource side, since the power architecture is fixed in the module. This means that even if a power upgrade may be required, the full - power architecture can be doubled or tripled by adding one or two identical modules respectively, which will not affect the physical behavior of the module. Therefore, the optimizer model remains effective without much NN training. In the hardware part, the power scheduling strategy provided by the resource optimizer 108 is controlled by the resource controller 116 to connect the power source to the main source bus.
[0040] Figure 3 is a flowchart showing the process of managing power according to one embodiment. Process 300 can be executed by Figure 1 resource level 101. Referring to Figure 3 , at block 301, the resource optimizer 108 receives the required total power from the central controller 109, where the required total power represents the total power to be consumed by the IT load 112, the thermal management system 113, and other subsystems 118. The subsystem 118 can include another set of loads similar to the load level 103, for example, as shown in FIGS. 6A and Figure 6BAs shown in. The intermediate level 102 will manage the power distribution to other subsystems. At block 302, the resource optimizer 108 determines the current status of power supplies 104 - 107, including which of the power supplies are available and their respective capacities, etc. Note that at block 302, these are also inputs to the optimizer, but they are not considered external variable inputs. The resource optimizer 108 can call the resource controller 116 to retrieve or determine the status of the power supplies. At block 303, the resource optimizer 108 calculates the optimized power required from different power supplies 104 - 107. In one embodiment, the resource optimizer 108 includes an NN model to determine the required optimized power based on the total power required and the status of power supplies 104 - 107. At block 304, the resource controller 116 receives the required optimized power from the resource optimizer 108 and configures and selects at least some of the power supplies 104 - 107 accordingly, which provide appropriate power to the intermediate level 102 at block 305. The resource level 101 is also referred to as level 1.
[0041] Back to Figure 1 , in one embodiment, the intermediate level 102 includes a power bus or interconnection 117 coupled between the output of the resource level 101 and the input of the load level 103 to transfer power from the resource level 101 to the load level 103. Additionally, the power bus 117 also provides power to other subsystems 118 other than the IT load 112 and the thermal management system 113. The intermediate level 102 also includes a subsystem load detector 119 (also referred to as an output load detector) and a central controller 109. The subsystem load detector 119 is configured to determine the power consumption of the subsystem 118 based on the workload of the subsystem 118. The central controller 109 is coupled to the subsystem load detector and the scheduling controller 110 to receive the subsystem power demand and the load power demand of the load level 103. In one embodiment, the central controller 109 includes an NN model to infer the total power required based on the power demands provided by the subsystem load detector 119 and the scheduling controller 110.
[0042] The intermediate level 102 mainly includes a power bus that connects the output of the resource level 101 to the input of the load level 103. There is a load detector or a system - to - system resource scheduling detector implemented. This is mainly for system - to - system power scheduling requirements. The output load detector 119 is used to provide the central controller 109 with the energy delivered to the load. The central controller 109 is an independent controller that takes inputs from both the load - level power demand from the scheduling controller 110 and other subsystem / system - to - system power demands from the output load detector 119, and sends the total power demand to the resource level 101 and monitors the output power from the resource level 101. The central controller 109 is configured to determine the total power required for the intermediate level 102 and the load level 103. The intermediate level 102 is also referred to as level 2.
[0043] This level is higher than the resource level and the load level of level 0 and level 1. The central controller 109 receives two power inputs from in - system controllers (e.g., the scheduling controller and the power flow optimizer) or system - to - system power controllers. There can be multiple Figure 1 systems 100, and they are interconnected. For example, the first subsystem is levels 101 and 103, while the second subsystem is another set of levels 101 and 103. The combination of these two subsystems is considered system - to - system and is connected by the intermediate level 102. The system - to - system controller receives power demands from its own load level 103 and the load levels 103 of other subsystems. In addition, the system - to - system controller receives the power output from its resource level 101 provided by the output load detector. It provides the required total power to level 1. This is in the case where the current central controller 109 does not respond to the power demands from other subsystems 118. A key design here is that the central controller 109 can add together the load demands from other subsystems 118 and then transmit the updated required total power to level 1. Another NN model is integrated with the central controller 109 for system - to - system power scheduling, such as in the case of power outages, blackouts, or system service or maintenance.
[0044] Figure 4 is a flowchart showing the process of operating the intermediate level according to one embodiment. The process 400 can be executed by the intermediate level 102. Refer to Figure 4, at block 401, the central controller 109 receives the load power demand from the scheduling controller 110. At block 402, the central controller 109 receives the power demands of other subsystems from the output load detector 119. The output load detector 119 receives the power demands of other subsystems, and the output load detector 119 provides how much power is provided by the resource level 103. At block 403, the NN model of the central controller 109 determines the total required power and the inter-system level scheduling strategy. At block 404, the central controller 109 outputs the total required power to the resource optimizer 108. For the intermediate level 102, it is connected to each of its own resources and power, and it also receives the required power from other subsystems.
[0045] Figure 5 Illustrates an overall three-level control hierarchy design and operation method according to one embodiment. The key connection is power or energy. The controller is integrated with the optimizer, and the optimizer is embedded with an NN model. The input and output of each controller are data representing power. The detailed power scheduling logic and principle do not affect between different levels. It can be seen that the changes in the input in each layer are isolated between layers and adopted within each layer. This is how decoupling is achieved while the entire system works as an organic whole. Note that each of the controllers and optimizers 108 to 111 and 116 can be implemented as a processor, microcontroller, ASIC, or FPGA, and each of them can include an NN model embedded therein.
[0046] In one embodiment, as Figure 1 shown, the data center system is one of the data center subsystems in a cluster. Figure 6A and Figure 6B illustrates a larger-scale system with multiple modules (or can be understood as multiple interconnected subsystems) according to one embodiment. Referring to Figure 6A and Figure 6B , the data center system includes subsystem 100A and subsystem 100B. Although only two subsystems are shown, more subsystems can be implemented. Each of subsystems 100A - 100B can represent the data center system 100 as shown in Figure 1 . Each of subsystems 100A - 100B includes their respective controllers (108A - B, 109A - B, 110A - B, 111A - B, and 116A - B), as described above with respect to Figure 1 .
[0047] Each of subsystems 100A - 100B is connected to Figure 1Same as shown. However, they are connected to the intermediate level 102 via the inter-system bus 150. This is why additional output power controllers and power detectors are used for each of the central controllers 109A - 109B. In this example, the central controllers 109A - 109B are referred to as level 3 controllers. Even if inter-system power scheduling is necessary and required. It does not affect the controllers or optimizers at level 0 (e.g., load level 103), and especially at level 1 (e.g., resource level 101). Since the only variable is the total power that has already been considered in the NN model in the optimizer, this is the benefit of decoupling because future upgrades may be needed by adding more and more subsystems to the inter-system bus 150. In some cases, each of the subsystems may not be the same. Even if the system is in a heterogeneous state, the individual optimizers can still work properly. In one embodiment, the central controller 1 and the central controller 2 are level 2 controllers at the intermediate level 102. The output power controller 1 and the output power controller 2 are level 3 controllers for communication and inter-system power scheduling. This means that the output power controllers 1 and 2 only communicate with the central controller.
[0048] Note that each of the output power controllers 1 and 2 can represent Figure 1 the output load detector 119. Each power controller is configured to receive power and / or load demands from another subsystem via the inter-system bus 150. In addition, each power controller can also provide the power and / or load demands of its own subsystem to another subsystem via the inter-system bus 150. Therefore, in this example, the output power controllers 1 and 2 communicate with each other within the intermediate level 102. One subsystem can provide power to another subsystem via the inter-system bus 150.
[0049] Figure 7 is a block diagram showing a multi-cluster design of a power system according to one embodiment. Generally, a data center can be hosted in one or more data center campuses. Each campus can include one or more data center buildings. Each building can host one or more clusters. Each cluster can include one or more data center subsystems, and each subsystem can include various modules or units. In each model, there are one or more level 0 and level 1. Level 0 and level 1 are connected by level 2. Level 2 is designed to connect level 1 and level 2. Cluster control is considered a level 3 controller, and the higher level, i.e., the central controller as shown, is a level 4 controller.
[0050] In this example, as Figure 7As shown, there are two clusters, and each cluster includes two subsystems. As shown in each of the modules, the number of loads or the number of resources can be different and can be upgraded, which does not affect any other system. Even within each module, upgrades or changes in level 1 and level 0 do not change the NN model in each of the optimizers because it is upgraded by repeating the same infrastructure.
[0051] As an example, referring to Figure 7 , subsystem 1_1 can be a GPU cluster, while subsystem 1_2 can be a general computing cluster. Any business model upgrade or business model change will be responded to only by dedicated modules. As another example, if one or more subsystems are added to cluster 1, even a brand-new subsystem 1_3 with a new IT and power / cooling system, it will have its own level 0 and level 1 optimizers and controllers within its module, and only the connection is through output controller 1_3. At the cluster level, i.e., at level 3 and level 2, the impact is minimal because they only communicate with the amount of power required and can be scheduled.
[0052] Therefore, the corresponding NN models in these layers can remain valid and do not require significant retraining. This design can be understood as a container-based solution, i.e., it includes the corresponding control policies and optimized NNs, so they can be reused for system expansion and upgrade through a decoupled design. The control hierarchy enables the diversification and variety of the system. In addition, it is beneficial to optimize the power efficiency and workload allocation design from the module to different layers of the entire campus.
[0053] Figure 8 is a flowchart showing a process for managing the power of a data center according to one embodiment. Process 800 can be executed by processing logic that can include hardware, software, or a combination thereof. Referring to Figure 8 , at block 801, the power flow optimizer determines the load power demand of the load portion based on the thermal data of the thermal management system and the load data of the electronic racks as loads. At block 802, the subsystem load detector determines the subsystem power demand of one or more subsystems. At block 803, the central controller in the middle portion determines the total power required based on the load power demand and the subsystem power demand. At block 804, the resource controller configures and selects at least some of the power supplies based on the total power required.
[0054] In the foregoing specification, embodiments of the present invention have been described with reference to specific exemplary embodiments of the present invention. Obviously, various modifications can be made to it without departing from the broader spirit and scope of the present invention as set forth in the appended claims. Accordingly, the specification and drawings are to be regarded as illustrative rather than restrictive.
Claims
1. A data center system, comprising: A load section, having a plurality of electronic racks, a thermal management system, and a power flow optimizer, wherein each of the electronic racks includes a plurality of servers, and each server contains one or more electronic devices, wherein the thermal management system is configured to provide cooling and / or heating to the electronic devices, and wherein the power flow optimizer is configured to determine the load power demand of the load section based on the thermal data of the thermal management system and the load data of the electronic racks; A resource section, having a plurality of power supplies for supplying power to the load section, wherein the resource section includes a resource controller that configures and selects at least some of the power supplies to supply power to the load section based on the load power demand provided by the power flow optimizer, wherein the power flow optimizer includes a power flow neural network model for predicting the amount of power required by the electronic racks and the thermal management system based on the thermal data and the load data to meet the thermal demand and data processing load demand of the load section; and An intermediate section, coupled between the resource section and the load section, wherein the intermediate section includes a power bus for distributing power from the resource section to the load section and other subsystems, The intermediate section further includes: a subsystem load detector, coupled to the other subsystems to determine the subsystem power demand; and a central controller, coupled to the power flow optimizer of the load section and the subsystem load detector, to determine the total power demand based on the subsystem power demand and the load power demand, wherein the resource controller utilizes the total power demand to configure and select at least some of the power supplies, and the central controller includes a neural network model for inferring the total power demand based on the load power demand and the subsystem power demand; The resource section further includes a resource optimizer for receiving the total power demand from the central controller to generate power supply configuration information, wherein the resource controller is configured to configure the power supplies based on the power supply configuration information, and the resource optimizer includes a neural network model for determining the power supply configuration information based on the total power demand; The data center system is the first data center subsystem among a plurality of data center subsystems, and wherein the power bus of the intermediate section is coupled to the power bus of the intermediate section of the second data center subsystem among the plurality of data center subsystems, and the neural network model of the central controller is further used to determine the inter-system power scheduling strategy of the data center subsystems.
2. The data center system according to claim 1, wherein, The load section further includes: One or more temperature sensors, disposed within each server to measure the temperature of the electronic devices; and A workload detector, configured to determine the workload of each of the servers, wherein the power flow neural network model infers the load power demand based on the temperature and the workload of each of the servers.
3. The data center system according to claim 2, wherein, The load portion further includes a power scheduling controller, which is coupled to the power flow optimizer to proportionally distribute the power received from the resource portion to the thermal management system and the server based on the load power demand received from the power flow optimizer.
4. The data center system according to claim 3, wherein, The load power demand includes information on how to schedule the power to the thermal management system and the server, and wherein the power scheduling controller is configured to output the total power required by the load portion to the central controller.
5. The data center system according to claim 1, wherein, The power supply configuration information includes information specifying the amount of power to be provided by each of the plurality of power supplies.
6. The data center system according to claim 1, wherein, The plurality of power supplies includes a utility power supply, a photovoltaic power supply, and a battery power supply.
7. The data center system according to claim 1, wherein, The central controller is shared by the first data center subsystem and the second data center subsystem.
8. The data center system according to claim 1, wherein, The plurality of data center subsystems is the first data center cluster in a plurality of data center clusters.
9. The data center system according to claim 8, wherein, Each of the data center clusters is controlled by a corresponding cluster controller, and wherein the central controller is shared by the plurality of data center clusters.
10. A method for managing a data center system, the method comprising: Using a power flow optimizer, determine the load power demand of the load portion having the thermal management system and a plurality of the electronic racks based on the thermal data of the thermal management system and the load data of the electronic racks, wherein each of the electronic racks includes a plurality of servers, and each server contains one or more electronic devices, and wherein the thermal management system is configured to provide cooling and / or heating to the electronic devices; Based on the load power demand provided by the power flow optimizer, configure and select at least some of the power supplies by a resource controller of a resource portion having a plurality of power supplies to provide power to the load portion, wherein the power flow optimizer includes a power flow neural network model for predicting the amount of power required by the electronic racks and the thermal management system based on the thermal data and the load data to meet the thermal demand and the data processing load demand of the load portion; The data center system further includes an intermediate portion, which is coupled between the resource portion and the load portion, and wherein the intermediate portion includes a power bus to distribute power from the resource portion to the load portion and other subsystems; The method for managing a data center system further includes: Using a subsystem load detector coupled to the other subsystems to determine the subsystem power demand; and Using a central controller coupled to the power flow optimizer and the subsystem load detector of the load portion, determine the total power demand based on the subsystem power demand and the load power demand, wherein the resource controller utilizes the total power demand to configure and select at least some of the power supplies, and the central controller includes a neural network model for inferring the total power demand based on the load power demand and the subsystem power demand; The resource section also uses a resource optimizer of the central controller connected to the intermediate section to receive the total power demand to generate power configuration information, wherein the resource controller is configured to configure the power supply based on the power configuration information, and the resource optimizer includes a neural network model for determining the power configuration information based on the total power demand; The data center system is the first data center subsystem among a plurality of data center subsystems, and wherein the power bus of the intermediate section is connected to the power bus of the intermediate section of the second data center subsystem among the plurality of data center subsystems, and the neural network model of the central controller is further used to determine the inter-system power scheduling strategy of the data center subsystem.
11. The method according to claim 10, further comprising: Measure the temperature of the electronic device using one or more temperature sensors provided in each server; And Use a workload detector to determine the workload of each of the servers, wherein the load power demand is inferred by the power flow neural network model based on the temperature and the workload of each of the servers.
12. The method according to claim 11, further comprising: Use a power scheduling controller connected to the power flow optimizer to proportionally allocate the power received from the resource section to the thermal management system and the servers based on the load power demand received from the power flow optimizer.
Citation Information
Patent Citations
Renewable energy based green data center load scheduling method and device
CN103377084A
Data center energy-saving scheduling method and system
CN109800066A