For processor-based node scheduling jobs to regulate coolant flow temperature
Patent Information
- Application Number
- CN202311114410.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2023-03-08
- Filing Date
- 2023-08-31
- Publication Date
- 2026-09-18
- Estimated Expiration
- 2043-08-31
AI Technical Summary
[0002]基于处理器的平台(例如,刀片服务器)的操作可能会产生相当大量的热能或废热,所述热能或废热如果不充分去除,将可能导致该平台的散热部件超过其热规范
Smart Images

Figure CN118625902B_ABST
Abstract
Description
Technical Field
[0001] This disclosure generally relates to scheduling jobs for processor-based nodes to regulate coolant flow temperature. Background Technology
[0002] The operation of processor-based platforms (such as blade servers) can generate a considerable amount of heat or waste heat, which, if not adequately removed, can cause the platform's heat dissipation components to exceed their thermal specifications. One way to remove waste heat from a computer platform is to circulate a stream of liquid coolant through a coolant flow plate (or "cold plate") located near the platform's heat dissipation components. Summary of the Invention
[0003] According to one aspect of this disclosure, a method for scheduling jobs is provided, comprising: transferring a coolant flow between an inlet and an outlet of a coolant subsystem associated with a cooling domain to remove heat from a plurality of processor-based nodes of the cooling domain, wherein the transfer of the coolant flow has associated predefined parameters, and wherein the predefined parameters include at least one or a combination of coolant density, a minimum volume of the coolant flow, and a coolant specific heat capacity; and adjusting the temperature of the coolant flow at the outlet, comprising: determining a minimum overall power consumption of the processor-based node based on the predefined parameters to maintain the temperature of the coolant flow at the outlet at or above a minimum threshold temperature; and scheduling jobs to be performed by the node based on the minimum overall power consumption.
[0004] According to another aspect of this disclosure, a system for scheduling jobs is provided, comprising: a rack-based computer subsystem including: a coolant supply manifold for delivering an inlet coolant flow; a coolant return manifold for delivering an outlet coolant flow; a plurality of cooling domains, wherein each of the plurality of cooling domains is associated with a coolant outlet temperature and includes: a flow path for delivering a coolant flow between the coolant supply manifold and the coolant return manifold, wherein the flow path includes an outlet coupled to the return manifold and providing the coolant outlet temperature; a processor-based node; and a power manager for scheduling jobs among the processor-based nodes of the plurality of cooling domains to regulate each coolant outlet temperature to maintain the coolant outlet temperature at or above a minimum temperature.
[0005] According to another aspect of this disclosure, a non-transitory storage medium is provided for storing machine-readable instructions that, when executed by a machine, cause the machine to: characterize a relationship between job execution performed by a plurality of processor-based nodes and a coolant outlet temperature of a coolant subsystem that removes heat from the plurality of processor-based nodes; determine a job schedule for the plurality of processor-based nodes based on the characterization and power profiles associated with the plurality of jobs to maintain the coolant outlet temperature above a minimum threshold temperature; and cause data representing the job schedule to be transmitted to a job manager of the plurality of processor-based nodes. Attached Figure Description
[0006] Figure 1 This is a schematic diagram of a computer system according to an example embodiment, which provides a coolant flow to a captured waste heat consumption system and schedules operations to regulate the temperature of the coolant flow.
[0007] Figure 2 This is a flowchart depicting a process, according to an example implementation, for determining an acceptable power consumption range for a task to be performed by a processor-based node in a cooling domain.
[0008] Figure 3 It is a flowchart depicting a process, according to an example implementation, for scheduling jobs to be executed by processor-based nodes in a cooling zone to regulate the temperature of the coolant flow provided by the cooling zone.
[0009] Figure 4 This is a sequence diagram illustrating the communication between components of a computer system according to an example embodiment, and the actions taken by the components to adjust the temperature of the coolant flow provided by the computer system to the captured waste heat consumption system.
[0010] Figure 5 This is a flowchart of a process for scheduling operations to maintain the temperature of the coolant stream at the outlet at or above a minimum temperature, according to an example implementation.
[0011] Figure 6 This is a block diagram of a system, according to an example implementation, for scheduling jobs among processor-based nodes in a cooling domain to maintain the coolant outlet temperature at or above a minimum temperature.
[0012] Figure 7 This is an illustration of a non-transitory storage medium storing machine-readable instructions according to an example embodiment, which, when executed by a machine, cause the machine to schedule jobs for processor-based nodes to maintain a minimum coolant outlet temperature at or above a minimum threshold temperature. Detailed Implementation
[0013] The following detailed description refers to the accompanying drawings. Where possible, the same reference numerals are used in the drawings and in the following description to refer to the same or similar parts. However, it should be clearly understood that these drawings are for illustrative and descriptive purposes only. While several examples are described in this document, modifications, adaptations, and other embodiments are possible. Therefore, the following detailed description does not limit the disclosed examples. Instead, the appropriate scope of the disclosed examples may be defined by the appended claims.
[0014] The terminology used herein is for the purpose of describing particular examples only and is not intended to be restrictive. As used herein, the singular forms “a,” “an,” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. As used herein, the term “plurality” is defined as two or more. As used herein, the term “another” is defined as at least a second or more. As used herein, unless otherwise indicated, the term “connected” is defined as a connection, whether directly connected without any intervening elements or indirectly connected by means of at least one intervening element. Two elements may be mechanically coupled, electrically coupled, or communicated through a communication channel, path, network, or system. As used herein, the term “and / or” refers to and covers any and all possible combinations of the associated listed items. It should also be understood that although the terms first, second, third, etc., may be used herein to describe various elements, these elements should not be limited by these terms, as these terms are only used to distinguish one element from another, unless otherwise stated or indicated by the context. As used herein, the term "includes" means including but not limited to, and the term "including" means including but not limited to. The term "based on" means at least partially based on.
[0015] The ever-increasing processing power of computer systems corresponds to a growing power consumption footprint. Due to factors such as rising energy prices and legislative frameworks promoting energy sustainability, improving the power efficiency of computer systems may be beneficial.
[0016] One method to improve the power efficiency of a computer system is to capture waste heat from the computer system and use the captured waste heat to drive another energy-consuming process. One way to capture waste heat from a computer system is to circulate a liquid coolant stream, such that the circulating coolant stream absorbs or captures waste heat generated by the system's heat dissipation components. The captured waste heat can then be transferred from the computer system to another system (referred to herein as the "captured waste heat consumption system"), which provides the energy-consuming process driven by the captured waste heat.
[0017] Energy-consuming processes can take many different forms. As an example, a waste heat capture system can be a heating, ventilation, and air conditioning (HVAC) system, such as an HVAC system in a building or a nearby building housing a computer system. For instance, an HVAC system may include a heat exchanger that transfers captured waste heat from the coolant flow from the computer system to an airflow to heat the building. As another example, an HVAC system may include an adsorption chiller or absorption chiller that uses the captured waste heat to drive a thermodynamic process for cooling water (e.g., for use in air conditioning).
[0018] As another example, the system that consumes the captured waste heat can be a power generation system. For instance, the power generation system may include a generator actuated by a turbine, and the captured waste heat may be used to generate steam to drive the turbine. As yet another example, the power generation system may use the captured waste heat to generate electricity in a thermoelectric process, such as a Seebeck process or a Peltier process.
[0019] Regardless of the specific application or use of the captured waste heat, a waste heat consumption system may have an associated specified minimum temperature for the coolant flow received by the system. Therefore, the coolant flow provided by the computer system (referred to herein as the “outgoing coolant flow” or “outgoing primary coolant flow”) may be subject to a minimum temperature specification, and the computer system may adjust the temperature of the outgoing coolant flow to meet that specification.
[0020] One way a computer system regulates the temperature of the outgoing coolant stream is by adjusting the volumetric flow rate of the coolant via one or more flow control devices (e.g., pumps and / or flow control valves). However, regulating the temperature of the outgoing coolant stream solely based on volumetric flow rate can be relatively inefficient in terms of energy.
[0021] According to the example implementations described herein, the computer system uses job scheduling to regulate the temperature of the outgoing coolant flow. According to some implementations, job scheduling-based temperature regulation can be used in conjunction with other mechanisms to regulate coolant temperature. For example, according to some implementations, the computer system can increase (e.g., maximize) the power consumed by nodes processing jobs in order to increase (e.g., maximize) the corresponding waste heat generated. As another example, according to some implementations, the computer system can use volumetric flow rate-based temperature regulation. According to yet another implementation, the computer system can use only job scheduling-based coolant temperature regulation.
[0022] More specifically, according to an example implementation, the computer system includes processor-based nodes (referred to herein as "nodes"), and the nodes are partitioned or grouped into cooling domains for thermal management purposes. As an example, the computer system may include one or more rack-based computer subsystems or "racks," and the nodes may be hosted by blade servers mounted in the racks(s). In this example, a particular blade server may host multiple nodes, and a cooling domain may correspond to the blade server and include the nodes hosted by the blade server. In another example, the nodes of a single blade server may be partitioned into multiple cooling domains. In other examples, a cooling domain may include nodes hosted by multiple blade servers.
[0023] Regardless of how nodes are divided between cooling domains, a liquid coolant flow (referred to herein as a "coolant flow") can circulate through the coolant flow plates or cold plates of the cooling domain to capture or absorb waste heat generated by the heat dissipation components of the cooling domain. In the example, the cooling domain includes blade servers, and the coolant flow can circulate through the cold plates of the blade servers to absorb waste heat generated due to the operation of the heat dissipation components of the blade servers, such as central processing units (CPUs) and / or graphics processing units (GPUs).
[0024] The cooling zone receives the incoming coolant flow, circulates it through multiple flow paths, and provides an outgoing coolant flow. According to an example embodiment, the coolant flow paths of the cooling zone may be part of a secondary coolant subsystem of the computer system. According to an example embodiment, the secondary coolant subsystem merges or combines individual cooling zone outgoing coolant flows to form a coolant outlet flow, which circulates through the secondary loop of a heat exchanger to transfer waste heat to the primary coolant subsystem of the computer system. The cooled coolant flow exits the secondary loop of the heat exchanger to provide the incoming coolant flow to the cooling zone.
[0025] According to the example implementation, the coolant flow of the primary coolant subsystem is received from the captured waste heat consumption system, circulated through the primary loop of the heat exchanger to absorb waste heat, and provided as an outgoing coolant flow (“primary outgoing coolant flow”) to the captured waste heat consumption system.
[0026] According to an example implementation, a computer system regulates the temperature of the primary outgoing coolant flow by adjusting job scheduling for nodes in the computer system's cooling domain to maintain the temperature at or above a minimum threshold temperature (e.g., a minimum coolant temperature specified for a system that captures waste heat). In this context, a "job" refers to a unit of work that can be assigned to a specific node. In the example, a set of jobs may be associated with a specific application and may be assigned to a set of nodes. A message passing interface (MPI) job is an example of a "job." In the example, a given job may be divided into a set of levels (e.g., MPI levels) or processes, and each level of a given job may be processed in parallel with other levels of said job. In the example, the computer system's job scheduler may schedule each level of a job to a specific set of nodes.
[0027] According to an example implementation, the computer system can set a target outgoing coolant flow temperature (or "target temperature") for a given cooling domain. The computer system can then schedule jobs for each node in each cooling domain in a manner that maintains the temperature of the outgoing coolant flow in the cooling domain at or above the target temperature of the cooling domain. Maintaining the outgoing coolant flow from the cooling domain at or above its respective target temperature results in the temperature of the composite primary outgoing coolant flow being at or above a minimum threshold temperature specified for the system consuming captured waste heat. Using job scheduling to regulate coolant flow temperature can be beneficial for several reasons, such as a reduced power footprint, faster transient response times, and adaptability to a wider range of processing loads.
[0028] As a more specific example Figure 1 A computer system 100 according to some embodiments is depicted. As described herein, the computer system 100 circulates a liquid coolant flow through various coolant flow paths to capture waste heat generated due to operations performed by a processor-based node 103 (hereinafter referred to as "node 103") of the computer system 100. The captured waste heat is transferred to an outgoing coolant flow (hereinafter referred to as "primary outgoing coolant flow 149"), which is delivered to a captured waste heat dissipation system 199. The captured waste heat dissipation system 199 has an associated specified minimum threshold temperature for the primary outgoing coolant outlet flow 149. As described herein, the computer system 100 schedules operations for node 103 in a manner that maintains the temperature of the primary outgoing coolant outlet flow 149 at or above the specified minimum threshold temperature.
[0029] Typically, computer system 100 may include various heat dissipation components, such as a central processing unit (CPU), a graphics processing unit (GPU), a network interface controller (NIC), power transistors, voltage regulators, and other heat dissipation components. Waste heat generated by at least some of these components (e.g., CPU and GPU) is related to job execution activities of the node 103 associated with these components.
[0030] According to an example embodiment, computer system 100 includes a primary coolant subsystem (or "facility-side coolant subsystem") that receives a primary incoming coolant stream 145 from a captured waste heat dissipation system 199. The primary coolant subsystem absorbs captured waste heat from a secondary coolant subsystem of computer system 100 and provides the primary outgoing coolant stream 149 to the captured waste heat dissipation system 199. According to an example embodiment, the secondary coolant subsystem circulates the coolant stream to absorb or capture waste heat from heat dissipation components of computer system 100, such that the captured waste heat can be transferred to the primary coolant subsystem. The coolant streams in the respective primary and secondary coolant subsystems are isolated from each other, and according to an example embodiment, a heat exchanger 150 transfers heat between the coolant subsystems.
[0031] Depending on the specific implementation, the coolant may be any of a variety of compositions, such as deionized water, distilled water, ethylene glycol, oil, a mixture of the aforementioned liquids, or another liquid or liquid mixture. Furthermore, depending on the specific implementation, the coolants in the first coolant subsystem and the second coolant subsystem may have the same composition or may have different compositions.
[0032] The primary coolant subsystem includes an inlet 144 that receives a primary incoming coolant stream 145 from a captured waste heat consumption system 199. The primary coolant subsystem further includes an outlet 148 that supplies the primary outgoing coolant stream 149 to the captured waste heat consumption system 199. The captured waste heat consumption system 199 transfers the heat energy from the received primary outgoing coolant stream 149 to drive any of several different energy-consuming processes, such as a thermodynamic process for cooling water, a process for heating air, a process for generating electricity, or another process. Therefore, the resulting primary incoming coolant stream 145 has a lower temperature than the primary outgoing coolant stream 149.
[0033] According to an example embodiment, coolant circulates in the primary coolant subsystem from inlet 144 to the primary loop inlet 163 of heat exchanger 150 and through the primary loop of heat exchanger 150 to absorb waste heat captured from the secondary coolant subsystem. The coolant exits the primary loop of heat exchanger 150 at a primary loop outlet 165 connected to outlet 148. According to some embodiments, the primary coolant subsystem may include a controllable valve 161 (e.g., a regulating valve) on the pipeline with inlet 144 and which can be used to regulate the flow volume of the primary coolant. According to an example embodiment, in addition to or alternatively thereto, the primary coolant subsystem may include one or more other components for regulating the flow volume, such as one or more pumps (not shown), a controllable valve coupled to the primary loop outlet 165, and other possible volume control components.
[0034] According to an example embodiment, the secondary coolant subsystem includes a rack-based coolant supply manifold 140 connected to an outlet 160 of the secondary cooling circuit of the heat exchanger 150. The secondary coolant subsystem further includes a coolant return manifold 168 connected to an inlet 164 of the secondary cooling circuit of the heat exchanger 150. Coolant circulates in the secondary coolant subsystem in a direction starting from the outlet 160 and passing through a port of the coolant supply manifold 140 to supply a coolant flow path circulating between the heat dissipation components of node 103 to absorb waste heat from said components. The coolant flow from the coolant flow path returns to a corresponding port of the coolant return manifold 168 to form a composite coolant flow received at the inlet 164 of the secondary cooling circuit of the heat exchanger 150. The coolant flow received at inlet 164 circulates through the secondary cooling circuit of the heat exchanger 150 to transfer waste heat to the coolant flow circulating in the primary coolant subsystem.
[0035] for Figure 1 In the specific example implementation depicted, computer system 100 includes a rack-based computer subsystem 101, wherein server trays 110 are mounted on a frame referred to as a "rack". The rack-based computer subsystem 101 may alternatively be referred to as a "rack-based computer system" or simply a "rack". Figure 1 In the particular example depicted, server tray 110 is arranged vertically in the rack. As another example, according to another embodiment, server tray 110 can be organized into horizontally arranged groups, wherein the groups are arranged vertically in the rack.
[0036] According to an example implementation, server tray 110 may include one or more blade servers 114. Although Figure 1Two blade servers 114 are depicted in each server tray, but according to another implementation, server tray 110 may have a single blade server 114 or may have more than two blade servers 114.
[0037] Blade server 114 is an example of a computer platform. According to another implementation, the computer system may have one or more computer platforms other than blade servers. As an example, a computer platform may be a rack-based server, network switch, gateway, tablet computer, edge processing subsystem, desktop computer, or any other electronic device including one or more processors (e.g., one or more CPU and / or GPU cores).
[0038] Figure 1 The route of the coolant flow path through a specific server tray 110-1 is specifically described. According to the example embodiment, other server trays 110 of the rack-based computer subsystem 101 may have similar coolant flow paths and characteristics. Turning now to the details of the specific coolant flow path of server tray 110-1, a coolant supply manifold 140 has an outlet port connected to a coolant supply line 132 via a disconnector 131. According to some embodiments, a flow control device 134 (e.g., a pump or controllable valve) may be connected between the coolant supply line 132 and the coolant supply line 130. The coolant supply line 130 has an outlet port connected to a corresponding coolant supply line 139 that provides coolant flow to the coolant flow path of the blade server 114.
[0039] According to an example embodiment, the coolant return manifold 168 has an inlet port that is connected to the coolant return line 170 via a disconnector 171. According to some embodiments, a flow control device 174 (e.g., a pump or controllable valve) may be connected between the coolant return line 170 and the coolant return line 180. The coolant return line 170 has an inlet port that is connected to receive coolant flow from the corresponding coolant return line 183 of the blade server 114 of the server tray 110-1.
[0040] Flow control devices 134 and 174, as well as similar flow control devices for other server trays 110, can be controlled to independently regulate the volume of coolant flow through server tray 110. In this way, according to some embodiments, one or two flow control devices 134 and 174 can be controlled for server tray 110-1 to regulate the volume of coolant flow through server tray 110-1. According to another embodiment, the secondary coolant subsystem may include a single flow control device (e.g., flow control device 134 or flow control device 174 for server tray 110-1) per server tray 110 to regulate the volume of coolant flow through server tray 110. As further described herein, computer system 100 may use multiple flow control devices per server tray 110 to regulate the volume of coolant flow through a blade server 114 of a given server tray 110 to regulate the temperature of the outflow coolant flow from server tray 110. According to yet another embodiment, the secondary coolant subsystem may not include any flow control device per server tray 110. For example, in these implementations, computer system 100 may use one or more flow control devices (e.g., multiple pumps and / or multiple controllable valves) to regulate the volume of coolant flow through manifolds 140 and 168, and may not have flow control devices for different server trays 110.
[0041] According to an example implementation, blade server 114 may include one or more coolant flow plates, referred to herein as "cold plate" 120. Typically, cold plate 120 may be physically located at or near the heat dissipation components of blade server 114 to absorb waste heat from the heat dissipation components and transfer the absorbed waste heat to the circulating coolant flow.
[0042] According to an example embodiment, the cold plate 120 is connected to a coolant flow path between the coolant supply pipe 130 and the coolant return pipe 180. As an example, according to some embodiments, the cold plate 120 may include one or more flow plates having serpentine flow channels for conveying coolant flow. According to an example embodiment, the cold plate 120 of a particular blade server 114 may form a serial chain-like coolant flow path. As an example, such as Figure 1As depicted, a specific coolant flow path for the blade server 114 may include four cold plates 120 arranged in a series chain. As an example, the first cold plate 120 in the series chain may include an inlet receiving an inlet coolant flow from an inlet coolant line 139. The outlet of the first cold plate 120 is connected to the inlet of the second cold plate 120 in the series chain coolant flow path. The outlet of the second cold plate 120 is connected to the inlet of the third cold plate 120 in the series chain coolant flow path. The outlet of the third cold plate 120 is connected to the inlet of the fourth cold plate 120 in the series chain coolant flow path. The outlet of the fourth cold plate 120 is connected to a coolant outlet line 183.
[0043] According to some embodiments, one or more sensors 137 may be inline coupled to and / or coupled to the coolant supply supply pipe 130. As an example, the sensors 137 may include temperature sensors that provide an indication (e.g., an analog signal or digital bits) representing the sensed temperature of the coolant flow in the pipe 130. According to some embodiments, one or more sensors 181 may be inline coupled to and / or coupled to the coolant return pipe 180. As an example, the sensors 181 may include temperature sensors that provide an indication (e.g., an analog signal or digital bits) representing the sensed temperature of the coolant flow in the pipe 180.
[0044] According to an example implementation, blade server 114 hosts or provides node 103. In the context used herein, a "node" refers to a logical or physical entity configured to execute machine-readable instructions to process jobs scheduled for that node. As an example of a "node," according to some implementations, a node may correspond to a hardware component running an instance of an operating system. For example, blade server 114 may have multiple multi-core CPUs, and a set of processing cores of a particular CPU may execute a particular operating system instance and be considered part of the same node. As a more specific example, according to some implementations, a particular blade server 114 may have two multi-core CPUs, each providing cores for two operating system instances and two corresponding nodes 104. As another example of a "node," according to some implementations, a node may be a single blade server 114 as a whole. As yet another example, a "node" may be multiple blade servers 114. As yet another example, a node may be a physical partition (e.g., a particular CPU or a set of CPU cores) of blade server 114, independent of an attached operating system instance.
[0045] For the discussion below, it is assumed that node 103 refers to a hardware component corresponding to a specific operating system instance, and may correspond to one or more processing cores of, for example, a specific blade server 114. Thus, for the discussion below, a specific blade server 114 may have multiple nodes 103 (e.g., two nodes 103 for each blade server 114).
[0046] According to an example implementation, the rack-based computing subsystem 101 is divided into cooling zones or cooling domains 111. Here, a "cooling domain" refers to a portion of the rack-based computing system 101 whose outgoing coolant flow temperature is regulated or controlled independently of other portions of the rack-based computing subsystem 101. As an example, cooling domain 111 may include one or more server trays 110. Figure 1 As depicted herein, according to some implementations, cooling domain 111 may include a single server tray 110. Therefore, according to an example implementation, cooling domain 111 may include multiple nodes 103.
[0047] As further described herein, in order to regulate the temperature of the coolant outlet flow from cooling zone 111, computer system 100 can control the volume of the coolant flow and control how jobs are scheduled for nodes 103 of the cooling zone. For example, for example server tray 110-1, according to an example implementation, computer system 100 can regulate flow control devices 134 and 174 and can control job scheduling for nodes 103 of server tray 110-1 to regulate the temperature of the coolant flow in coolant return pipe 180.
[0048] According to another embodiment, the cooling zone 111 can be different from... Figure 1 The nodes 103 depicted in the diagram correspond to the grouping. As an example, according to another implementation, cooling domain 111 may include multiple server trays 110. As yet another example, according to another implementation, a given server tray 110 may be divided among multiple cooling domains 111.
[0049] According to some embodiments, the computer system 100 may have cooling domains in addition to the server blade-based cooling domains. For example, according to some embodiments, the computer system 100 may include cooling domains that include one or more power supplies for a rack-based computer subsystem 101. As another example, according to some embodiments, the computer system 100 may include cooling domains for power distribution units (PDUs) of the rack-based computer subsystem 101. As another example, according to some embodiments, the computer system 100 may include cooling domains for different portions of the PDUs. The computer system 100 may regulate the temperature of the outlet coolant flow from these cooling domains through coolant volume regulation and / or other means (e.g., power regulation).
[0050] According to some embodiments, computer system 100 includes a system power manager 186 configured to schedule jobs in a manner that regulates the temperature of primary outgoing coolant flow 149, such that the temperature is maintained at or above a specified minimum threshold temperature. According to some embodiments, system power manager 186 may consider various control parameters to determine the target outgoing coolant flow temperature of a given cooling zone 111. These control parameters may include, for example, the temperature of primary incoming coolant flow 145 (e.g., the temperature provided by temperature sensor 147) and / or the temperature of primary outgoing coolant flow 149 (e.g., the temperature provided by temperature sensor 143). As another example, control parameters may include the CPU utilization per node 103 for each cooling zone 111. As another example, control parameters may include the currently sensed outgoing coolant flow temperature of each cooling zone 111. As another example, control parameters may include the sensed or calculated current coolant flow volume of each cooling zone 111. As another example, control parameters may include the temperature of the incoming coolant flow to each cooling zone 111.
[0051] As an example, according to some implementations, the system power manager 186 may set the same target outgoing coolant flow temperature for each cooling zone 111. As another example, according to yet another implementation, the system power manager 186 may independently determine the target temperature for each cooling zone 111, which may result in varying target temperatures. Regardless of the method used, according to the example implementation, the system power manager 186 sets the target temperature for each cooling zone 111 and uses job scheduling to adjust the corresponding actual temperature to its corresponding target temperature or above, such that the temperature of the primary outgoing coolant flow 149 is maintained at or above the minimum threshold temperature specified for the captured waste heat dissipation system 199.
[0052] The following describes how the system power manager 186 can schedule jobs for nodes 103 of a specific cooling domain 111 to maintain the temperature of the outflow coolant flow of the cooling domain at a target temperature (referred to herein as "T"). EXIT (Temperature) or above. According to the example implementation, the system power manager 186 determines the minimum total power level consumption (hereinafter referred to as "P") of node 103 of cooling domain 111. NODE_MIN (Power consumption) to maintain the temperature of the coolant flow from cooling zone 111 at T EXIT Temperature or above. According to the example implementation, the system power manager 186 can then schedule a batch of jobs for node 103 to maintain the overall power consumption of node 103 at P based on the job's power profile. NODE_MIN Power consumption or above. Here, the "power profile" of a job refers to an indication or representation of the power that node 103 is expected to consume when processing the job.
[0053] After scheduling these jobs, the system manager 186 can then pass the corresponding job schedule to the job power manager 189, which, according to an example implementation, optimizes the execution of each scheduled job to improve power consumption. For example, according to some implementations, job power limits or power caps may exist to push the application from which the job originates to a more energy-efficient (instructions per watt) operating point. To improve (e.g., maximize) the power consumption generated by the node executing the job (and thus increase the corresponding waste heat), the job power manager 189 can remove any such power caps for the job. As a specific example, according to some implementations, to improve the corresponding power consumption, the job power manager 189 can disable the dynamic power management features of the node 103 executing the job.
[0054] As another example of increasing the power consumed by processing jobs, the job power manager 189 can enable a turbo mode for node 103, which allows node 103 to exceed its rated specifications if node 103 has the power and thermal headroom to handle such processing. For example, a given node 103 can be rated to operate at a maximum frequency of 3 GHz, and in turbo mode, node 103 can operate at frequencies exceeding the specified maximum frequency (e.g., 3.2 GHz, 3.4 GHz, or 3.6 GHz) for short periods.
[0055] In addition to scheduling operations and increasing power consumption to regulate the temperature of the coolant flow out of cooling zone 111, according to an example embodiment, computer system 100 can regulate the volume of the coolant flow. More specifically, according to some embodiments, system power manager 186 can send a target temperature of cooling zone 111 to cooling manager 188 of computer system 100. Cooling manager 188 can perform actions such as regulating flow control devices 134 and 174 to maintain the temperature of the coolant flow out of cooling zone 111 at or above the target temperature. In another example, cooling manager 188 can regulate the opening of valve 161 to maintain the temperature of primary coolant flow 149 at or above a minimum threshold temperature. As another example, cooling manager 188 can regulate the pump (not shown) of the primary coolant flow subsystem to maintain the temperature of primary coolant flow 149 at or above the minimum threshold temperature. As another example, cooling manager 188 can regulate the operation of controllable valves in cooling zone 111 or the operation of controllable valves in the secondary coolant subsystem.
[0056] According to some implementations, the system power manager 186, cooling manager 188, and / or job power manager 189 may be entities formed by one or more processors 192 executing machine-executable instructions 191 (or "software"). As an example, the processor 192 may be one or more CPU processing cores located on one or more servers 184. The servers(s) 184 may include one or more servers as part of a rack-based computer subsystem 101 and / or may include one or more servers remotely configured relative to the rack-based computer subsystem 101 (e.g., configured in a geographic location other than the data center where the rack-based subsystem 101 is configured).
[0057] As an example, according to an example implementation, a particular server 184 may be a chassis management controller that is part of a rack-based subsystem 101. As another example, according to some implementations, a particular server 184 may be a cloud-based server. As yet another example, according to some implementations, node 103 may be part of a high-performance computing (HPC) cluster, and system power manager 186 may be executed by processor 192 of one or more management nodes of the cluster; and cooling manager 188 and job power manager 189 may be executed by processor 192 of the chassis management controller of rack-based subsystem 101.
[0058] According to another implementation, one or more of the system power manager 186, cooling manager 188, and / or operating power manager 189 may be formed by dedicated hardware (e.g., logic gates) that performs one or more functions without executing machine-executable instructions. In this way, depending on the specific implementation, the hardware may be an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a complex programmable logic device (CPLD), etc.
[0059] Usually, such as Figure 1 As depicted by reference numeral 181 in the accompanying drawings, servers 184 may be coupled to server blades 114 via rack interconnect and / or network structure 182. As an example, the rack interconnect may include a backplane, backplane connectors, bus interfaces, network cables, one or more network switches, and possibly other and / or different rack interconnect components. Network structure 182 may be one or more types of communication networks, such as (as an example) a Fibre Channel network, a Compute Fast Link (CXL) architecture, a private management network, a local area network (LAN), a wide area network (WAN), a global network (e.g., the Internet), a wireless network, or any combination thereof.
[0060] Now turning to the more specific details of job scheduling, according to the example implementation, the system power manager 186 determines the P for each cooling zone 111. NODE MIN Minimum power consumption. P NODE MIN The minimum power consumption level is determined by the consumption of node 103 in a given cooling domain 111 to maintain the temperature of the outflow coolant flow from the cooling domain at T. EXIT Minimum total power at or above the target temperature. System power manager 186 can schedule jobs for nodes 103 of a given cooling domain 111 such that the total power consumption of the job meets or exceeds the P value determined for the cooling domain 111. NODE MIN Minimum power consumption.
[0061] The following describes the relationship between cooling system parameters and node power consumption for a given cooling domain 111 according to an example embodiment. The heat power captured or absorbed by the coolant flow of cooling domain 111 (referred to herein as "Q-absorbed heat power") can be described as follows: (Equation 1) In Equation 1, "V" represents the volumetric flow rate of the coolant. "D" represents the volumetric flow rate of the coolant. W "C" indicates the density of the coolant at the operating temperature and pressure (e.g., the density of water). W "T" indicates the specific heat capacity of the coolant at the operating temperature and pressure. in"T" indicates the inlet temperature of the coolant flow (e.g., the temperature of the coolant flow at or near the outlet port of the coolant supply manifold 140), and "T" out "Indicates the outlet temperature of the coolant flow (e.g., the temperature of the coolant flow at or near the inlet port of the coolant return manifold 168).
[0062] As can be seen from Equation 1, the thermal power Q changes directly with the volume V of the coolant flow. Therefore, an increase in the volume V flow rate leads to more waste heat being captured by the coolant flow, and vice versa. Furthermore, as can also be observed from Equation 1, an increase in volume V leads to a decrease in T... OUT The coolant outlet temperature decreases, and the reduction in V volume will cause T OUT The coolant outlet temperature increases.
[0063] The heat power absorbed by Q is directly related to the total power consumption of node 103 and is proportional to the effectiveness of the cooling system. In this context, "cooling system effectiveness" refers to the waste heat captured or absorbed by the coolant flow. According to the example implementation, the cooling manager 188 can be constrained to regulate the volume V of the coolant flow within a certain volumetric flow rate range, the lower boundary of which is referred to herein as "V". MIN "Minimum volume". According to the example implementation, P NODE MIN Given V MIN In the case of minimum volume, node 103 will T OUT The coolant outlet temperature is maintained at T EXIT Minimum total power consumption at the target temperature.
[0064] According to the example implementation, the system power manager 186 can determine the P of a given cooling zone 111 as follows: NODE MIN Minimum power consumption: (Equation 2) Here, "|" represents the absolute value operator. Because the T value of cooling zone 111... IN The coolant inlet temperature can vary over time, so according to some implementations, the system power manager 186 can recalculate P from time to time. NODE MIN Minimum overall power consumption. For example, whenever the system power manager 186 schedules a job for node 103 of cooling domain 111, the system power manager 186 can calculate the P of a specific cooling domain 111. NODE MIN Minimum overall power consumption. As another example, whenever the system power manager 186 initiates a process for scheduling a batch of jobs for node 103 of cooling domain 111, the system power manager 186 can calculate the P of a specific cooling domain 111. NODE MIN Minimum overall power consumption. As another example, the system power manager 186 can calculate P based on the schedule.NODE MIN Minimum overall power consumption. As another example, the system power manager 186 can calculate the P of a given cooling domain 111 in response to events other than those related to the scheduling of a job or a batch of jobs (e.g., blade server 114 or a group of blade servers 114 powering on, exiting a reset, or powering off). NODE MIN Minimum overall power consumption.
[0065] Figure 2 An example process 200 for determining the acceptable power consumption range of a task, according to an example implementation, is described. As an example, Figure 1 The job power manager 186 can perform process 200 before scheduling jobs in a batch of jobs for a specific cooling domain 111 node 103, and evaluate candidate jobs in the batch of jobs using an acceptable power consumption range. In this way, as further described herein, according to an example implementation, the job power manager 186 can determine whether the expected power consumption of a given candidate job (as indicated by the job's power profile) falls within an acceptable power consumption range, and based on this determination, determine whether the candidate job should be scheduled as part of a batch of jobs.
[0066] refer to Figure 2 According to some implementations, process 200 includes accessing (box 204) representing one or more configuration parameters, V MIN Minimum coolant volume, cooling zone T IN Coolant inlet temperature and T CASE Temperature data. According to the example implementation, T CASE The temperature is the minimum of the maximum permissible surface temperatures of the components (e.g., semiconductor packages) in cooling zone 111. CASE Temperature effect on the cooling zone T EXIT The choice of temperature imposes a constraint. In this way, T CASE Temperature and V MIN Minimum coolant volume and V MAX The maximum coolant level is defined together with T. EXIT The upper and lower boundaries of the temperature. As an example, the configuration parameters for the cooling domain could be cooling system effectiveness parameters, D... W Coolant density and C W Coolant specific heat capacity. According to box 206, process 200 includes determining T. EXIT The upper and lower boundaries of the temperature, and within these boundaries, the cooling zone 111 is selected at T. EXIT temperature.
[0067] According to block 208, process 200 includes determining the maximum node power consumption. The maximum node power consumption of node 103 may be based on thermal standards of one or more hardware components of node 103. According to some implementations, the maximum node power consumption may be the minimum of the maximum node power consumptions of node 103 in cooling domain 111.
[0068] According to box 216, process 200 includes determining P. NODE_MIN Minimum overall power consumption. The minimum power consumption for each job defines the lower bound of the acceptable node power consumption range. As an example, for all nodes 103 in cooling domain 111, the acceptable node power consumption range can be the same, and the lower bound of the range is P. NODE_MIN The minimum total power consumption is divided by the number of nodes 103 in cooling domain 111, and the upper boundary of the range is the maximum node power consumption determined in block 208.
[0069] As another example, the acceptable node power consumption range can be specific for each node 103. For instance, the acceptable node power consumption range for a given node 103 can have an upper boundary corresponding to the maximum power consumption limit of node 103, and a lower boundary of the power range is P. NODE_MIN Minimum total power consumption divided by the number of nodes 103.
[0070] As another example, process 200 can set different lower boundaries for individual acceptable node power consumption ranges, where the sum of the lower boundaries is equal to or close to P. NODE_MIN Minimum overall power consumption.
[0071] Figure 3 An example process 300 according to some implementations is described, which can be used to schedule a batch of jobs for nodes in a cooling domain. As an example, process 300 can be initiated in response to a node in the cooling domain indicating an idle state or the availability of jobs to be executed. As an example, process 300 can be... Figure 1 The system power manager 186 is executed.
[0072] refer to Figure 3 According to some implementations, process 300 includes accessing data representing a power profile of a candidate job and an acceptable range of node power consumption for the node, according to block 308. The data may also represent power profiles of jobs being executed or planned to be executed on other nodes that are cooling data. Process 300 then includes generating a batch job containing multiple jobs for node execution, according to subprocess 310.
[0073] According to some implementations, the batch job generation sub-process 310 includes evaluating candidate jobs. More specifically, sub-process 310 may first include determining (decision block 312) whether the predicted power consumption (as indicated by the associated power profile) of the next candidate job is within the boundary of the acceptable node power consumption range. If not, then according to decision block 316, it is determined whether a specific strategy allows the addition of a candidate job, even if the candidate job is not within the acceptable node power consumption range. According to some implementations, the specific strategy may, for example, specify that if the predicted power consumption is outside the acceptable node power consumption range, the candidate job is not scheduled. Therefore, according to block 320, the candidate job is bypassed.
[0074] According to another implementation, the policy may allow a predetermined deviation from a strict application of the acceptable node power consumption range. For example, the policy may specify that a candidate job can be scheduled if the predicted power consumption is within a specified percentage of the lower boundary of the acceptable node power consumption range. As another example, process 300 may include predicting the total power consumption of all nodes that have been scheduled in the batch of jobs so far, and allowing a certain deviation about the lower boundary based on the predicted total power consumption.
[0075] If a candidate job is to be added to the batch job (e.g., the "Yes" branch of decision box 312 or the "Yes" branch of decision box 316), then according to box 324, process 300 includes adding the candidate job to the batch job. Next, according to decision box 328, process 300 includes determining whether another candidate job should be added to the batch job. This determination may be based on any of several factors, such as the target size of the batch job, the expected duration of the batch job's processing time, the predicted total power consumption of jobs already scheduled in the batch job, the existence of compatible candidate jobs (e.g., jobs from the same application or tier), or one or more other factors. If another job is to be added, then as... Figure 3 As depicted, control returns to decision box 312. Otherwise, according to an example implementation, process 300 includes an indication (box 322) that batch job processing is ready for execution. According to an example implementation, this indication may include transferring the batch jobs to a job power manager, such as... Figure 1 The power manager 189.
[0076] Figure 4 Sequence diagram 400 is depicted, illustrating the communication between system power manager 186, operation power manager 189, and cooling manager 188 according to an exemplary embodiment, and the actions performed thereby. Sequence diagram 400 appears in accordance with timeline 401.
[0077] refer to Figure 4Sequence diagram 400 first depicts (according to timeline 401) the scheduling job of system power manager 186. More specifically, scheduling includes system power manager 186 reading data at 420 representing cooling configuration, system cooling domain definitions, and power constraints. Furthermore, as indicated at 424, system power manager 186 reads Vt representing each domain. MIN Minimum volumetric flow rate and T EXIT Data on the minimum outlet coolant temperature. Furthermore, as depicted at 428, the system power manager 186 can communicate with the cooling manager 188 to read data representing the inlet temperature of the cooling zone. Based on this information, the system power manager 186 can then determine the P for each cooling zone at 432. NODE_MIN Minimum power consumption level and acceptable node power consumption range. As indicated at 436, the system power manager 186 can then base its power management on the power range, P... NODE_MIN The system uses power consumption and job power profiles to schedule jobs. This scheduling may include the system power manager 186 sending messages to the job power manager 189.
[0078] Figure 4 Actions taken by the job power manager 189 are depicted, including initiating a job for each job as depicted at 450 and communicating with the cooling manager 188 (as depicted at 454) to set a target outlet temperature X for the cooling domain.
[0079] The job power manager 189 can then optimize job execution to achieve maximum power consumption, as depicted at 464. At the end of each job, the job power manager 189 can then transmit a job completion instruction to the system power manager 186 (as depicted at 480) and a job completion instruction to the cooling manager 188 (at 470). In response to the job completion instruction from the job power manager 189, according to an example implementation, the system power manager 186 can add power margin from idle nodes to nodes still performing jobs, as depicted at 484. Furthermore, in response to job completion, as depicted at 474, the cooling manager can adjust the coolant flow volume of the idle nodes (e.g., by adjusting the flow control device to reduce the coolant flow volume), and adjust the coolant flow volume at the facility site (e.g., by adjusting...). Figure 1 Valve 161 is used to reduce the volume of the primary coolant flow.
[0080] refer to Figure 5According to an example implementation, process 500 includes transferring (box 504) a coolant flow between the inlet and outlet of a coolant subsystem associated with the cooling domain to remove heat from multiple processor-based nodes of the cooling domain. According to an example implementation, a node may be a node hosted by one or more computer platforms, such as blade servers. According to some implementations, a node may be associated with an operating system instance. According to some implementations, the cooling domain may include one or more computer platforms, such as one or more blade servers. According to some implementations, the cooling domain may include a portion of a computer platform (e.g., a portion of a blade server) or multiple computer platforms (e.g., multiple blade servers). According to an example implementation, the cooling domain may correspond to a heat dissipation component included in a sub-section of a rack-based computing system.
[0081] The delivery of the coolant flow has associated predefined parameters. According to an example embodiment, the predefined parameters may include coolant density. According to some embodiments, the predefined parameters may include a minimum volume of coolant flow. According to an example embodiment, the predefined parameters may include a coolant-specific specific heat capacity.
[0082] Process 500 includes regulating (block 506) the temperature of the coolant flow at the outlet. According to an example implementation, regulation includes determining (block 508) a minimum overall power consumption of the processor-based node based on predefined parameters to maintain the temperature of the coolant flow at the outlet at or above a minimum temperature threshold, and, according to block 512, scheduling jobs to be executed by the node based on the minimum power consumption. According to an example implementation, a job refers to a unit of work to be executed by a specific node. According to an example implementation, an acceptable power range for each job can be determined. According to an example implementation, the lower boundary of the acceptable power range can be determined based on the minimum overall power consumption. According to an example implementation, the upper boundary of the acceptable power range can be determined based on thermal limitation criteria for hardware components associated with the node.
[0083] According to an example implementation, the temperature of the coolant flow at the outlet can be further regulated by increasing (e.g., maximizing) the node's power consumption during processing or performing a task. As an example, node power consumption can be increased by removing power caps or restrictions associated with the node's energy-saving operation. As another example, node power consumption can be increased by placing the node in turbine mode.
[0084] According to example embodiments, the temperature of the coolant flow at the outlet can be further regulated by adjusting the coolant flow volume. According to some embodiments, adjusting the coolant flow volume may include adjusting the speed of one or more flow control devices associated with the cooling zone. According to some embodiments, adjusting the coolant flow volume may include adjusting the operation of a facility-side valve (or a primary coolant subsystem valve). According to example embodiments, the coolant flow from the outlet may be combined with coolant flows from other cooling zones to form a composite cooling flow containing captured waste heat from the cooling zones. According to some embodiments, the coolant flow from the cooling zones may be circulated through a heat exchanger to transfer the captured waste heat from the cooling zones to the coolant flow in the primary coolant subsystem. According to some embodiments, the primary coolant subsystem may circulate the coolant flow to provide an outlet coolant flow to a system consuming captured waste heat.
[0085] According to an example implementation, the coolant flow volume of the cooling zone can be limited in response to one or more nodes of the cooling zone becoming idle. According to an example implementation, the coolant flow volume on the facility side can be limited in response to one or more nodes becoming idle.
[0086] refer to Figure 6 According to an example implementation, system 600 includes a rack-based computer subsystem 610 and a power manager 650. The rack-based computer subsystem 610 includes a coolant supply manifold for conveying an inlet coolant flow; a coolant return manifold 618 for conveying an outlet coolant flow 622; and a plurality of cooling zones 620. Each cooling zone 620 is associated with a coolant outlet temperature and includes a flow path 630 to convey a coolant flow between the coolant supply manifold 614 and the coolant return manifold 618. The cooling zone 620 further includes a processor-based node 640. The flow path 630 may include one or more coolant flow plates or cold plates.
[0087] Rack-based computer subsystem 610 may include blade servers associated with nodes. Node 640 may be associated with an operating system instance. Cooling domain 620 may include one or more blade servers.
[0088] According to an example implementation, the power manager 650 may be provided by a chassis management controller of a rack-based subsystem 601, which executes machine-executable instructions or software. According to another example implementation, the power manager 650 may be provided by a server remotely configured relative to the rack-based subsystem 601.
[0089] The power manager 650 schedules jobs among processor-based nodes 640 to regulate each coolant outlet temperature 622 to maintain the coolant outlet temperature 622 at or above the minimum temperature.
[0090] According to an example implementation, system 600 may further include a job power manager that manages the execution of jobs by node 640. The job power manager may be implemented by one or more processors (or in software) that execute machine-readable instructions, by hardware that does not execute machine-readable instructions, or by a combination of such software and hardware. The job power manager may increase (e.g., maximize) the power consumption of node 640 when processing or executing jobs. As an example, the job power manager may increase node power consumption by removing power caps or limits associated with power-saving operation of node 640. As another example, the job power manager may increase node power consumption by placing node 640 in turbo mode.
[0091] According to an example implementation, system 600 may further include a cooling manager that manages the coolant flow volume to regulate the coolant temperature. The cooling manager may be implemented by one or more processors (or in software) executing machine-readable instructions, by hardware that does not execute machine-readable instructions, or by a combination of such software and hardware. According to some implementations, the cooling manager may regulate the coolant flow volume by adjusting one or more flow control devices associated with cooling zone 620. According to some implementations, the cooling manager may regulate the coolant flow volume by adjusting the operation of facility-side valves. According to an example implementation, the cooling manager may reduce the coolant flow volume of cooling zone 620 in response to one or more nodes 640 of cooling zone 620 becoming idle. According to an example implementation, the cooling manager may reduce the facility-side coolant flow volume in response to one or more nodes 640 becoming idle.
[0092] refer to Figure 7 According to an example embodiment, the non-transitory storage medium 700 stores machine-readable instructions 710, which, when executed by a machine, cause the machine to characterize the relationship between job execution on a processor-based node and the coolant temperature of a coolant subsystem for removing heat energy from the processor-based node. According to an example embodiment, the machine may include a chassis management controller, and the instructions 710 may be executed by one or more processors (e.g., CPU cores) of the chassis management controller.
[0093] According to an example implementation, a node may be a node hosted by one or more computer platforms, such as blade servers. According to some implementations, a node may be associated with an operating system instance. According to some implementations, a cooling domain may include one or more computer platforms, such as one or more blade servers. According to some implementations, a cooling domain may include a portion of a computer platform (e.g., a portion of a blade server) or multiple computer platforms (e.g., multiple blade servers). According to an example implementation, a cooling domain may correspond to a heat dissipation component included in a sub-section of a rack-based computing system.
[0094] When executed by the machine, instruction 710 further causes the machine to determine job scheduling for multiple processor-based nodes based on characteristics and power profiles associated with multiple jobs, in order to maintain the coolant outlet temperature at or above a minimum threshold temperature. Instruction 710, when executed by the machine, further causes the machine to transmit data representing the job scheduling to the job manager of the processor-based node. According to an example implementation, the job manager can manage job execution in a manner that improves (e.g., maximizes) node power consumption.
[0095] According to an example implementation, scheduling includes: determining candidate job profiles based on minimum overall power consumption; and selecting a given candidate job to be executed based on the power profile of the given candidate job. Among the potential advantages is improved power efficiency of the computer system.
[0096] According to the example implementation, determining the candidate job profile includes determining the allowable range of power consumption levels. Among the potential advantages is the potential to improve the power efficiency of the computer system.
[0097] According to an example implementation, scheduling includes: determining power profiles for a plurality of candidate jobs; and selecting a subset of candidate jobs based on the power profiles. Among the potential advantages is improved power efficiency of the computer system.
[0098] According to an example implementation, scheduling includes: determining a second power consumption less than the minimum power consumption based on the minimum power consumption; and scheduling jobs based on the second power consumption. The method further includes regulating the delivery of the coolant flow to maintain the temperature of the coolant flow at the outlet at or above a minimum threshold temperature during job execution by the node. Among the potential advantages is improved power efficiency of the computer system.
[0099] According to an example implementation, the method further includes adjusting the delivery of coolant flow during operation performed by the node to maintain the coolant temperature at the output at or above a minimum threshold temperature. Among the potential advantages is improved power efficiency of the computer system.
[0100] According to an example implementation, regulating the delivery of coolant flow includes at least one of the following: regulating the operation of a coolant pump associated with the plurality of processor-based nodes, wherein the coolant pump is disposed on a server blade tray containing the plurality of processor-based nodes; regulating a flow control valve upstream of a supply manifold that supplies coolant flow to the coolant subsystem and at least one other coolant subsystem associated with a plurality of additional processor-based nodes; or regulating a flow control valve downstream of a return manifold that supplies coolant flow. Among the potential advantages is improved power efficiency of the computer system.
[0101] According to an example implementation, the method further includes managing the execution of tasks to regulate the temperature of the coolant at the outlet. Among the potential advantages is improved power efficiency of the computer system.
[0102] According to an example implementation, regulating coolant delivery includes reducing the volume of the coolant flow in response to the completion of tasks by the plurality of processor-based nodes. Among the potential advantages is improved power efficiency of the computer system.
[0103] Although this disclosure has been described with respect to a limited number of embodiments, those skilled in the art who benefit from this disclosure will recognize many modifications and variations. The appended claims are intended to cover all such modifications and variations.
Claims
1. A method for scheduling jobs, comprising: A coolant flow is transferred between the inlet and outlet of a coolant subsystem associated with a cooling domain to remove heat from multiple processor-based nodes of the cooling domain, wherein the transfer of the coolant flow has associated predefined parameters, and wherein said predefined parameters include at least one or a combination of coolant density, minimum volume of the coolant flow, and coolant specific heat capacity; and Adjusting the temperature of the coolant flow at the outlet includes: The minimum overall power consumption of the processor-based node is determined based on the predefined parameters to maintain the temperature of the coolant flow at the outlet at or above a minimum threshold temperature; and The jobs to be executed by the nodes are scheduled based on the minimum overall power consumption.
2. The method of claim 1, wherein, The scheduling includes: Candidate job profiles are determined based on the minimum overall power consumption; and A given candidate job is selected from a plurality of candidate jobs to be executed by at least one of the nodes, the selection being based on the power profile of the given candidate job.
3. The method as described in claim 2, wherein, Determining the candidate job profile includes determining the allowable range of power consumption levels.
4. The method of claim 1, wherein, The scheduling includes: Determine the power profiles of multiple candidate jobs; A subset of candidate jobs is selected from the plurality of candidate jobs based on the power profile.
5. The method of claim 1, wherein, The scheduling includes determining a second power consumption less than the minimum total power consumption based on the minimum total power consumption, and scheduling the job based on the second power consumption, the method further including: The delivery of the coolant is regulated so that the temperature of the coolant flow at the outlet is maintained at or above the minimum threshold temperature during the operation performed by the node.
6. The method of claim 1, further comprising: During the operation performed by the node, the volume of the coolant flow is adjusted to maintain the temperature of the coolant flow at the outlet at or above the minimum threshold temperature.
7. The method of claim 6, wherein, The delivery of the coolant flow regulation includes at least one of the following: Regulate the operation of coolant pumps associated with the plurality of processor-based nodes, wherein the coolant pumps are disposed on a server blade tray containing the plurality of processor-based nodes; Adjust the flow control valve upstream of the supply manifold, which supplies the coolant flow to the coolant subsystem and at least one other coolant subsystem associated with a plurality of additional processor-based nodes; Adjust the flow control valve downstream of the return manifold that provides the coolant flow, or The combination of the above items.
8. The method of claim 1, further comprising: The execution of the operation is managed to regulate the temperature of the coolant flow at the outlet.
9. The method of claim 1, wherein, The regulation of the coolant flow delivery includes reducing the volume of the coolant flow in response to the completion of the task by the plurality of processor-based nodes.
10. A system for scheduling jobs, comprising: A rack-based computer subsystem, comprising: A coolant supply manifold for conveying an inlet coolant flow; A coolant return manifold, the coolant return manifold being used to deliver an outlet coolant flow; Multiple cooling zones, wherein each of the multiple cooling zones is associated with a coolant outlet temperature and includes: A flow path for conveying a coolant flow between the coolant supply manifold and the coolant return manifold, wherein the flow path includes an outlet coupled to the return manifold and providing the coolant outlet temperature; and Processor-based nodes; and A power manager is configured to schedule jobs among the processor-based nodes in the plurality of cooling domains to adjust the coolant outlet temperature to maintain the coolant outlet temperature at or above a minimum temperature.
11. The system of claim 10, wherein, For a given cooling domain among the plurality of cooling domains, the power management controller is further configured to: Determine the minimum overall power consumption of the processor-based nodes in the given cooling domain to maintain the coolant outlet temperature of the given cooling domain at the minimum temperature; and The processor-based node scheduling job for the given cooling domain is based on the minimum overall power consumption.
12. The system of claim 10, wherein, A given cooling zone among the plurality of cooling zones includes a pump connected to a flow path of the given cooling zone, and the system further includes: A cooling manager is configured to regulate the operation of the pump during operations performed by processor-based nodes in the cooling domain to maintain the coolant outlet temperature at or above the minimum outlet temperature.
13. The system of claim 12, wherein, The coolant management controller is used to further regulate the operation of the pump in response to the processor-based node of the cooling domain transitioning to an idle state.
14. The system of claim 12, wherein, A given cooling domain among the plurality of cooling domains includes a blade server tray, and the processor-based node is mounted to the blade server tray.
15. The system of claim 12, further comprising: A job manager that receives data from a power manager that schedules a given job in the jobs, wherein the job manager optimizes the job to improve power consumption.
16. A non-transitory storage medium for storing machine-readable instructions, which, when executed by a machine, cause the machine to perform the following operations: Characterizes the relationship between job execution performed by multiple processor-based nodes and the coolant outlet temperature of a coolant subsystem that removes heat from the multiple processor-based nodes; Based on the characterization and power profiles associated with multiple jobs, job scheduling for the multiple processor-based nodes is determined to maintain the coolant outlet temperature above a minimum threshold temperature. as well as This causes the data representing the job scheduling to be transmitted to the job manager of the plurality of processor-based nodes.
17. The storage medium of claim 16, wherein, When executed by the machine, the instructions further enable the machine to characterize the relationship based on the minimum coolant flow volume, the effectiveness of the coolant subsystem, and the highest processor node temperature.
18. The storage medium of claim 16, wherein, When executed by the machine, the instructions further cause the machine to perform the following operations: Candidate job profiles are determined based on the aforementioned characteristics; as well as A given candidate job is selected from a plurality of candidate jobs to be executed by at least one of the plurality of processor-based nodes, the selection being based on the power profile of the given candidate job.
19. The storage medium of claim 18, wherein: The characterization includes the total power consumption of the plurality of processor-based nodes; and The candidate job profile includes a target power consumption based on the average of the total power consumption.
20. The storage medium of claim 18, wherein: The characterization includes the total power consumption of the plurality of processor-based nodes; and The candidate job profile includes a target power range based on the total power consumption.
Citation Information
Patent Citations
Energy-saving scheduling method based on airflow organization for a data center
CN109871268A
Temperature control apparatus of fuel cell system and method of operating fuel cell system
CN115377450A