RACK-LEVEL MODULAR SERVER AND STORAGE FRAMEWORK
The modular information handling system framework addresses data center cooling and maintenance challenges by implementing shared cooling and power management, enhancing efficiency and reducing costs through centralized control and resource sharing.
Patent Information
- Application Number
- DE102011085335
- Authority / Receiving Office
- DE · DE
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2010-11-04
- Filing Date
- 2011-10-27
- Publication Date
- 2025-07-03
- Estimated Expiration
- 2031-10-27
AI Technical Summary
Data centers face challenges with high cooling costs and increased maintenance due to excess heat generated by servers, and the failure of active components necessitates time-consuming servicing, which increases system costs.
A modular information handling system framework with shared cooling, power, and management systems, utilizing a rack-level design with a shared fan module, power distribution unit, and centralized domain controller to manage and monitor multiple servers efficiently, minimizing post-installation maintenance and optimizing resource sharing.
The system minimizes maintenance costs and optimizes system efficiency by allowing servers to share cooling and power resources, reducing cooling costs and minimizing downtime due to component failures.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present invention relates generally to information processing systems and more particularly to a rack-level modular server and storage framework. BACKGROUND
[0002] As the value and use of information continues to grow, individuals and businesses are looking for additional ways to process and store information. One option available to users is information handling systems. Generally, an information handling system processes, translates, stores, and / or transmits information or data for business, personal, or other purposes, enabling users to benefit from the value of the information. Because technology and information handling needs and requirements vary between different users or applications, information handling systems can vary in what information is processed, how the information is processed, how much information is processed, stored, or transmitted, and how quickly and effectively the information can be processed, stored, or transmitted.The differences between information handling systems allow information handling systems to be general-purpose or configured for a specific user or application, such as financial transaction processing, airline reservations, corporate data storage, or global communications. Information handling systems may also include a variety of hardware and software components that can be configured to process, store, and transmit information and may include one or more computer systems, data storage systems, and network systems.
[0003] An information processing system, such as a server system, may be placed in a rack (rack-mounted enclosure). A rack may house multiple server systems, and multiple racks are typically placed in a space known as a data center or server room. A typical server room comprises rows of racks. One challenge with data centers is the heat generated by multiple servers in the data center. Excess heat results in high cooling costs for a data center and can lead to degraded performance of the rack's or data center's computer systems. In addition, servers often contain active components. Once a server is installed in a rack, the failure of an active component of the server can necessitate servicing, which increases system costs and can be time-consuming.
[0004] It is desirable to effectively manage and monitor servers located within a data center to minimize post-installation maintenance costs associated with the servers. Additionally, it is desirable to achieve optimal system efficiency by allowing servers to share system resources, such as fans required for server cooling and server power distribution units.
[0005] US 2006 / 0 041 783 A1 describes a container for use in connection with a data storage device and a housing of a computer peripheral enclosure having a housing that has at least one bay. Each bay has a rectangular cross-section and a depth. Each bay also has two guide rails. The container includes a U-shaped support and two guide rails. The U-shaped support further has two side walls. Each guide rail is mechanically connected to the side walls of the U-shaped support. The data storage device has two side walls, each having two mounting holes. The container slides freely but with a snug fit onto the guide rails into the bay. The U-shaped support has two inwardly directed projections that are aligned with the mounting holes of the data storage devices so that the storage devices can be securely arranged in the U-shaped support.
[0006] US Patent No. 6,025,989 A describes a modular node assembly for a rack-mounted multiprocessor computer, the node assembly comprising a logic subrack and a removable subrack. The logic subrack contains logic cards, memory cards, service processor cards, central processing unit (CPU) cards, and input / output cards connected by cables connecting input / output and processors. The removable subracks contain power supply modules for the node, a node monitoring card, hard disk drives, and fans. The removable subrack can be removed from the logic subrack without moving or disturbing the logic subrack.One fan of a fan pair draws air through the power supply in the removable subrack, which has a relatively higher cooling requirement, and blows the air into the logic subrack over the logic modules, which have a relatively lower cooling requirement. The other fan draws air over the hard disk drives and the node monitoring card, which have a relatively low cooling requirement, and blows air into the logic subrack over the logic modules, which have a relatively higher cooling requirement.
[0007] US 2006 / 0 007 651 A1 describes a blade server system comprising a chassis, a centrally located circuit board, a server blade, and a display panel. The server blade is electrically connected to the centrally located circuit board, while the display panel is electrically connected to the centrally located circuit board and is slidably mounted within the chassis for displaying the operating behavior of the server blade. SUMMARY
[0008] The present disclosure relates generally to information handling systems and more particularly to a rack-level modular server and storage framework.
[0009] In an exemplary embodiment, the present invention is directed to a modular information handling system framework. The modular information handling system may include a rack containing at least one chassis; a sled placed in the chassis; wherein the sled contains at least one information handling system; a fan placed in the chassis to cool the information handling system; a fan controller communicatively coupled to the fan; wherein the fan controller manages operation of the fan; a node controller associated with the sled; wherein the node controller manages operation of the sled; a power module to provide electrical power to the information handling system; a power module controller to manage the power module;and a primary domain controller communicatively coupled to the fan controller, the node controller, and the power module; wherein the primary domain controller manages the operation of the at least one fan controller, the node controller, and the power module.
[0010] In another exemplary embodiment, the present invention is directed to a modular rack system. The modular rack system may include a plurality of chassis placed in one or more racks; a plurality of sleds placed in each chassis; each sled comprising an information handling system; a shared fan module for cooling the plurality of sleds in each chassis; a shared power module for providing electrical power to one or more sleds in one or more chassis; and a shared management module for managing the operation of the plurality of chassis.
[0011] The methods and systems disclosed herein therefore provide for efficient management and monitoring of information handling systems that may be located in a data center, thereby minimizing post-installation maintenance costs. Furthermore, the methods and systems of the present application optimize system efficiency by allowing two or more information handling systems to share system resources such as power supplies and fans. Further technical advantages will become apparent to those skilled in the art upon review of the following description, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] A more complete understanding of the present embodiments and their advantages may be obtained with reference to the following description taken in conjunction with the accompanying drawings, in which like reference numerals indicate similar features and wherein: The Fig. 1 is a pictorial view of a modular rack system in accordance with an exemplary embodiment of the present invention. The Fig. 2 is a pictorial view of a chassis in accordance with an exemplary embodiment of the present invention. The Fig. 3 is a perspective view of the subrack of the Fig. 2. The Fig. 4 is a close-up view of a modular rack system in accordance with an exemplary embodiment of the present invention. The Fig. 5 is a block diagram of a system management framework for the modular rack system in accordance with an exemplary embodiment of the present invention. The Fig. Figure 6 is a block diagram of the software stack for the fan controller of the Fig. 5 in accordance with an exemplary embodiment of the present invention. The Fig. 7 is a block diagram of a software stack for the node controller of the Fig. 5 in accordance with an exemplary embodiment of the present invention. The Fig. Figure 8 is a block diagram of the software architecture of the domain controller of the Fig. 5 in accordance with an exemplary embodiment of the present invention. The Fig. 9 is a shared power system in accordance with an exemplary embodiment of the present invention. Fig. 10 illustrates the connection of a chassis to a power distribution unit in accordance with an exemplary embodiment of the present invention. The Fig. 11 illustrates a management device in accordance with an exemplary embodiment of the present invention.
[0013] Although embodiments of this disclosure have been illustrated and described and defined with reference to exemplary embodiments of the disclosure, such references do not imply any limitation of the disclosure, and no such limitation is intended to be inferred. The disclosed subject matter is capable of considerable modification, alteration, or equivalents in form and function, as will be apparent to those skilled in the art and who will appreciate the benefits of this disclosure. The illustrated and described embodiments of this disclosure are merely examples and are not exhaustive of the scope of the disclosure. DETAILED DESCRIPTION
[0014] For purposes of this disclosure, an information handling system may include any means or arrangement of means operable to compute, classify, process, transmit, receive, retrieve, originate, switch, store, display, communicate, verify, extract, reproduce, manipulate, or use any form of information, intelligence, or data for business, scientific, management, or other purposes. An information handling system may be, for example, a personal computer, a network storage device, or any other suitable device, and may vary in size, shape, performance, functionality, and price.The information processing system may include random access memory (RAM), one or more processing resources such as a central processing unit (CPU) or hardware or software control logic, ROM, and / or other types of non-volatile memory.
[0015] Additional components of the information handling system may include one or more hard disk drives, one or more network ports for communication with external devices, as well as various input and output (I / O) devices such as a keyboard, a mouse, or a video display. The information handling system may further include one or more buses operable to transmit communications between various hardware components.
[0016] An information processing system can be housed in a rack, such as the Fig. 1, servers and / or data storage devices in a data center may be arranged in a rack 102. As will be appreciated by those skilled in the art, the server may consist of at least a motherboard, a CPU, and memory. A data center may include one or more racks 102, depending on the user's system requirements. The rack 102 may include one or more chassis 104. A chassis 104, in accordance with an exemplary embodiment of the present invention, is a modular component that facilitates sharing critical server components among many servers. In an exemplary embodiment, the chassis 104 may be a 4U chassis. Each chassis 104 is conveniently free of active components internally or on its backplane to minimize service requirements after its installation.
[0017] As detailed in the Fig. 2 and Fig. 4, the chassis 104 may include sleds 106. The sled 106 may include one or more servers 107. The rack 102 may include one or more computational sleds, storage sleds, or a combination thereof. As will be apparent to those skilled in the art based on this disclosure, although the computational sleds and / or the data storage sleds are arranged vertically in the exemplary embodiment, they may also be arranged horizontally. As shown in the Fig. 2 and Fig. 3, an exemplary embodiment of the chassis 104 may support up to ten vertical compute-centric sleds, five double-wide sleds comprising twelve or more hard disk drives, or a hybrid arrangement including a combination of compute-centric and storage sleds. As those skilled in the art will appreciate, the present invention is not limited to any specific number or configuration of sleds 106 in a chassis 104. In an exemplary embodiment, in the case of horizontal sleds, four full-width 1U sleds supporting a dense server in each 1U, including four socket systems, may be used. As will be apparent to those skilled in the art, with the benefit of this disclosure, other arrangements of the compute-centric and storage sleds may be used depending on system requirements and / or preferences.
[0018] The subrack 104 may further include a shared fan module in a cooling path 110 at the rear of the subrack 104. In an exemplary embodiment, the shared fan module has a 4U resolution. In one embodiment, three fans 108 in the shared fan module may be used to cool the sleds 106 in the subrack 104. However, more or fewer fans may be used in the shared fan module depending on system performance and system requirements. The fans 108 may be managed by a fan controller 508, the operation of which is described in more detail below in connection with the Fig. 5 and Fig. 6. As will be appreciated by those skilled in the art having the benefit of this disclosure, the fan controller 508 may be hot swapped from the back of the rack 102 without removing power from the rack 102.
[0019] In addition, each subrack 104 may receive electrical power via cables coming from the Power Distribution Unit (“PDU”), which is described in more detail below in connection with the Fig. 5, Fig. 9 and Fig. 10. As discussed below, the PDU 902 may include one or more power supply units ("PSUs") 904. Therefore, the electrical power generated by all PSUs 904 is distributed among all chassis 104 connected to the PDU 902. In turn, each chassis 104 will then distribute the electrical power it receives to the individual sleds 106 contained within that chassis.
[0020] Between the cooling section 110 and the sleds 106 are backplanes 112. The chassis 104 may include a power and management backplane that distributes electrical power to each of the sleds 106. The power and management backplane may further carry high-speed network signals (such as Ethernet) and low-speed network signals (such as the System Management Bus). In one embodiment, the system may further include an optional storage backplane that allows a compute-centric sled to access one or more storage sleds in the same chassis 104 via SATA / SAS signals. The storage backplane connectors may be connected to the compute-centric backplane connectors via SATA / SAS patch cables.
[0021] As in Fig. 5, the system and methods disclosed herein disclose a framework for shared cooling, shared power, and shared management of the sleds 106. Each sled 106 within the chassis 104 may be treated as a node that is centrally managed via a network, such as an Ethernet-based management network. A domain controller 514 may provide an access point for a user to manage the system. The operating modes of the domain controller 514 are discussed in more detail below. Accordingly, as in the Fig. 5, each sled may further include a node controller 502 that provides system management and monitoring capabilities. The node controller 502 may manage each of the servers 107 in a sled 106. The node controller 502 manages the power of the servers 107, turns on the LED lights, reads temperatures, forwards data from the serial console, etc. It is also responsible for providing the appropriate power rails to supply electrical power to the components of the server 107. An end user 518 does not interact directly with the node controller 502. Instead, the node controller 502 communicates with the domain controller 514 and other devices on the management network.The term "management network" refers to an internal network, not accessible to an end user, that enables various controllers within the network to communicate with each other, and specifically with the domain controller 514. In an exemplary embodiment, the management network may be an Ethernet 10 / 100 network. The term "domain," as used herein, refers to a logical concept indicating a set of sleds 106, fans 108, power supplies 510, 116, and other devices arranged within a rack 102 or distributed across a set of racks that may be managed by the same domain controller 514.
[0022] The network may connect a consolidation switch 506 within each chassis 104 to a central management switch 516 at a central management domain controller 514, which provides a single (redundant) access point for a user through an interface such as, for example, a command line interface, a Simple Network Management Protocol, or a data center manageability interface. The domain controller 514 enables an end user to manage and monitor a domain. For example, the domain controller 514 may manage all sleds 506, fans 108, and power units 510, 116 in one or more chassis 104. The domain controller 514 communicates with the subordinate controllers using the management network.As discussed herein, the term "low-level controller" refers to node controllers, fan controllers, and power controllers that provide functionality but are not directly accessible to end user 518. Domain controller 514 may include a robust software stack to provide end user 518 with many different ways to manage the system. In one embodiment, the system may include two domain controllers 514, 524. If primary domain controller 514 fails, an automatic failover process may occur, and secondary controller 524 may take over and become the primary domain controller.
[0023] Under normal operating conditions, the primary controller 514 may have a connection to the secondary domain controller 524 through the management network. As will be appreciated by those skilled in the art, with the benefits of this disclosure, any number of suitable methods may be used to provide this connection. In one exemplary embodiment, the connection may be a TCP connection. The primary domain controller 514 may send an "I'm alive" message to the secondary domain controller 524 through the TCP connection every few seconds. The primary domain controller 514 may also send important updates to the secondary domain controller 524, such as registration messages, alarms, etc., through the TCP connection.The secondary domain controller 524 operates in a loop that checks the timestamp of the last "I'm live" message received from the primary domain controller 514.
[0024] If the second domain controller 524 goes offline or otherwise becomes inoperable while the primary domain controller 514 is operational, the primary domain controller 514 will detect that the secondary domain controller 524 cannot be reached (the TCP connection is down). An alarm will then be generated. The primary domain controller 514 will then attempt to re-establish the TCP connection (pause a few seconds between attempts). If a successful TCP connection is established with the second domain controller 524, an event will be generated to notify the system that the error has been resolved.
[0025] If the primary domain controller 514 goes offline or otherwise becomes inoperative while the secondary domain controller 524 is operational, the secondary domain controller 524 will no longer receive an "I'm live" message. If the secondary domain controller 524 does not receive an "I'm live" message after a predetermined time has elapsed, it will detect that the primary domain controller 514 has become inoperative. In response, the secondary domain controller 524 may generate an alarm to the system and / or change its operating mode and become the primary domain controller for the system. The subordinate controllers may not immediately notice the change in the domain controller. As a result, some "old" sensor data packets may be lost while the transition occurs.However, up-to-date sensor data will be cached once the secondary domain controller takes over. Similarly, as a result of the failure of the primary domain controller 514, a user interface (e.g., a command line interface or a web service) to the primary controller may be interrupted. However, after a few seconds, when the transition occurs, a new connection attempt will be successful, and the user can repeat the commands. Next, the new primary controller will attempt to establish a TCP session with a new secondary domain controller.
[0026] The Fig. 11 illustrates a management device 1100 in accordance with an embodiment of the present invention. The management device 1100 may include up to two domain controllers 514, 524 and an Ethernet switch 516. A system in accordance with an embodiment of the present invention may include up to one management device 1100 per rack 102. In examples where a user desires system redundancy, the management device 1100 may include both a primary domain controller 514 and a secondary domain controller 524. In contrast, if a user does not require redundancy, the management device 1100 may include only a primary domain controller 514.Furthermore, multi-rack systems that do not require redundancy may include one rack 102 with one management device 1100 having a domain controller and an Ethernet switch, and the other management device 1100 having only one switch. A management device 1100 according to an embodiment of the present invention may include a rotary switch that can be set with the position of the rack 102, so that each rack 102 can be assigned a different number. This is part of the position information that each device in the system has. For example, a sled 106 will have an assigned rack / chassis / sled identification.
[0027] Now returning to the Fig. 5, to manage and monitor the sleds 106, the node controller 502 associated with a sled 106 may have a physical connection to sensors, the motherboard, or expansion boards that provide management of the multiple servers 107 in a sled 106. The node controller 502 may run on a microprocessor. In one embodiment, the microprocessor may include a small embedded operating system. One of the primary responsibilities of the node controller 502 may be to provide and manage electrical power to the sleds 106, including the motherboards, hard disk drives, and / or storage sleds associated with a sled 106. Power management commands may be provided to the node controller from a panel on the chassis 104, through the management network, and / or a baseboard management controller ("BMC").Node controller 502 may be responsible for collecting data and alarming. Node controller 502 may then periodically send data to domain controller 514 and / or send notifications to domain controller 514 when events of interest occur. In addition, node controller 502 may send sensor data to the fan controller of chassis 508. For example, in addition to sending chassis sensor data, such as temperature sensor data, to domain controller 514 for storage, node controller 502 may also send sensor data to fan controller 508 of chassis 104. Fan controller 508 may then use the sensor data to control the speed of fans 108.
[0028] The fan controller 508 may include software to control and monitor the speed and status of the fans 108 and to notify the domain controller 514 of critical events of the fans 108. The fan controller 508 may communicate with a domain controller 514 via the management network. The fan controller 508 may receive temperature data from all node controllers 502 located in the same chassis 104 to regulate the fan speeds to meet the thermal requirements of the system. The fan controller 508 may include a master configuration file that its various components must read upon startup and that can be overwritten by the domain controller 514. In particular, the parameters that control the behavior of the fan controller 508, such as the pulling frequency, the default debug levels, etc., must be read from the configuration file and can be overwritten by the domain controller 514 for testing and tuning purposes.
[0029] Now the Fig. Turning to Figure 6, a block diagram of the components of fan controller 508 is shown. The fan controller may include a network abstraction layer 602 and a hardware abstraction layer 604. The connections between the services in fan controller 508 may be implementation-dependent and are largely influenced by the underlying hardware. For example, the connections may simply be calls to load libraries, shared memory, or interprocess communication frameworks such as queues. Network abstraction layer 602 enables fan controller 508 to send and receive messages from the network without being concerned with the underlying network protocol in use. Fan controller 508 may include an identification service 606 that determines the physical location of fan controller 508.In particular, the first task of fan controller 508 may be to identify itself. Since fan controller 508 is located within a chassis 104 of rack 102, this identification may be based on the physical location of the fan controller. Using hardware abstraction layer 604, fan controller 508 will determine the chassis number assigned to it. The chassis number will be known as a location string and may be assigned to a host name of fan controller 508. As those skilled in the art will appreciate, with the benefit of this disclosure, a static address may be assigned to fan controller 508 based on the physical location of fan controller 508. Once an IP address is assigned to fan controller 508, any other service in fan controller 508 must be restarted.The fan controller 508 will then validate that the address is unique within the network. If the assigned address is not unique, the fan controller 508 will send an error to the log service 614 and attempt to obtain an address from a reserved address pool.
[0030] In one embodiment, a dynamic fan controller 608 may be provided to regulate the speed of the fans 108. Sensors (not shown) may be placed in the system. The dynamic fan controller 608 may periodically receive sensor readings from one or more sensors of the chassis 104 and may dynamically adjust the speed of the fans 108 using a PID controller algorithm fed with sensor data from the sleds 106 in the chassis 104 and other environmental sensors placed in front of the chassis 104. For example, the dynamic fan controller 608 may receive the following sensor data from each sled 106: output ambient temperature (based on temperature samples from the node controller 502); CPU temperature (from the BMC); DIMM temperature (from the BMC); and sled power consumption.Additionally, the fan controller 608 may periodically receive sensor readings from the chassis 504, such as the ambient temperature. For each of the sensor readings, there will be a discrete PID controller in the dynamic fan controller 608. As will be appreciated by those skilled in the art, with the benefit of this disclosure, the PID controllers may control the fan speed based on the one or more variables received from the sensors used in the system. If there is a sensor failure, if the fan controller 508 fails, or if the dynamic fan control 608 otherwise fails and cannot be recovered, the fans will be commanded to operate at maximum speed.
[0031] Since the operation of such a feedback control system is known to those skilled in the art, it will not be discussed in detail herein. If one of the fans 108 of the fan module fails, the fan controller 508 will command the remaining fans to operate at maximum speed. In an exemplary embodiment, in the event of a firmware failure, the fans 108 may be caused to operate at maximum speed while the fan controller 508 is restarted.
[0032] The notification service 610 of the fan controller 508 can send messages from the fan controller 508 to the domain controller 514 and other recipients. The messages can include data updates or events of interest (e.g., fan failures). The first task of the notification service 610 is to notify the domain controller 514 that the fan controller 508 is ready. Furthermore, after the initial "registration," the notification service can forward messages received from other components of the fan controller 508 to the domain controller 514 and other devices (e.g., to the dynamic fan control 608). The fan controller 508 can further include a command listener service 612 that receives messages or commands from the domain controller 514 through a connection-oriented session that has been previously created.The command listener 612 can queue incoming requests and process them sequentially. The maximum size of the queue can be read from the configuration file. As a result, the methods executed by the command listener 612 to perform management and monitoring operations do not need to be thread-safe, although using a more thread-safe method is recommended. Although only a single connection from the domain controller 514 is required under normal operating conditions, it is desirable to have the ability to allow more than one connection to the queue for debugging purposes, so that a test client can send commands to the fan controller 508 even when connected to a domain controller 514.
[0033] The fan controller 508 may further include a log service 614 that receives messages from other components of the fan controller 508 and stores them in a physical medium, which may be a permanent location (e.g., an EEPROM). The log service 614 may rotate the log entries in the physical medium so that it is never full and that the most recent messages remain available. The maximum size of the log entry depends on the available hardware resources and may be part of the configuration file. For example, in one embodiment, the number of messages in the log service 614 may be 500, while in another embodiment, it may be 20.
[0034] Additionally, the fan controller 508 may include a monitoring service 616 that obtains the last reading for each sensor of interest (e.g., a speed sensor) and fires events of interest (e.g., a fan fault) to the notification service 610 when the sensor data falls outside a predetermined acceptable range. Furthermore, the monitoring service 616 may send periodic dynamic data updates to the domain controller 514 via the notification service 610. In one embodiment, the monitoring service 616 may continuously poll the hardware abstraction layer 604 for each sensor at a predetermined frequency and may store a predetermined number of sensor readings in memory.The stored sensor readings can then be used to calculate an average value for the specific sensor, which is reported when the monitoring service 616 is queried about a sensor. The number of sensor data points to be stored and the sampling rate can be set in the configuration file.
[0035] In one embodiment, the monitoring service 616 of the fan controller 508 may use received sensor data and compare it to three operating ranges to determine whether the fans 108 are operating in the normal range, the warning range, or the alarm range. Each time a sensor enters one of these ranges, the monitoring service 616 of the fan controller 508 may fire an event to the notification service 610, which will notify the end user 518 via the domain controller 514.
[0036] The scopes for each category can be set in the configuration file by the domain controller 514.
[0037] Finally, the fan controller 508 may include a heartbeat signal 618, which is a low-end device that polls the fan controller 508 at a predetermined frequency and resets the fans to operate at full speed if it does not receive a response from the fan controller 508.
[0038] To create flexible and maintainable code, the fan controller services may be arranged so that they do not interact directly with the hardware. Instead, the fan controller 508 may include a hardware abstraction layer 604 that acts as an interface between the services and the hardware 620. For example, if the command listener service 612 receives a command to turn off a fan 108, the command listener 612 may send a request to the hardware abstraction layer 604, which knows the physical medium and protocol for performing the task. As will be appreciated by those skilled in the art, with the benefits of this disclosure, a fan controller 508 may manage a number of hardware devices 62, including, but not limited to, a fan PWM 620a, a fan tachometer 620b, an EEPROM / Flash 620c, and a "fail no harm" controller 620d.
[0039] Now returning to Fig. 5, node controller 502 may perform management and monitoring operations requested by domain controller 504, such as server power management. Node controller 502 may communicate with domain controller 514 using the management network.
[0040] The domain controller 514 may include a master configuration file that the various components of the node controller 502 must read to start up. This configuration file can be overwritten by the domain controller 514. Accordingly, the parameters that control the performance of the node controller 502, such as the polling frequency, the default debug levels, etc., must be read from the master configuration file and can be overwritten for testing or tuning purposes. The presence of the master configuration file removes hard coding from the code and enables performance with minimal modifications during system testing.Additionally, a copy of the original configuration file can be retained in the system to enable a "rollback" where the original configuration file is written to the node controller and the system is rebooted.
[0041] Now the Fig. Turning to Figure 7, a block diagram of some example components of node controller 502 is shown. Node controller 502 may include a number of well-defined services and may make use of a network abstraction layer 702 and a hardware abstraction layer 704, which provide flexibility and portability to the software. The connection between the services may be implementation-dependent and may have an impact on the underlying hardware. In some embodiments, this may be a simple method for calling loaded libraries, shared memory, or inter-process communication frameworks such as queues.
[0042] The network abstraction layer 702 may enable software to send or receive messages from the network without concern for the underlying network protocol being used. One of the first tasks of the node controller 502 is to identify itself to the system. In one embodiment, the node controller 502 may identify itself to the system by specifying its physical location, which may consist of a specific rack 102, a chassis 104, and a sled 106. Accordingly, one of the first components to start the system may be the identification service 706, which determines the physical location of the node controller 502. Using the hardware abstraction layer 704, the node controller 502 will determine the chassis number and the node number within the chassis 104 where it is located.A static address can be assigned to the location of the specific node controller 502. Once an IP address is assigned, all other services of node controller 502 must be restarted. Node controller 502 can then ensure that the assigned address is unique for the network. If the address is not unique, node controller 512 can write an error and attempt to obtain an address from the reserved pool. The identification process can be performed frequently (for example, every 10 seconds), and if the location changes, it should re-register.
[0043] The notification service 708 may send messages from the node controller 502 to the domain controller 514 and other recipients. These messages may include data updates (e.g., sensor data) and / or events of interest (e.g., a change in state and errors). The first task of the notification service 708 is to notify the domain controller 514 that the node controller 502 is ready and to "register" the node controller 502 with the domain controller 514. If the initial attempt to register the node controller 502 was unsuccessful, the notification service 708 may then wait a predetermined time and then try again until a connection is established. Additionally, the notification service 708 may forward messages from other services and / or other modules to the node controller 502 through the management network.In one embodiment, the notification service 708 may send messages to the domain controller 514 at predetermined time intervals to detect the unlikely event that both the primary domain controller 514 and the secondary domain controller 524 (as discussed in more detail below) are both down. Once registration has been completed, the notification service 708 may read sensor and other dynamic data from the hardware managed by the node controller 502; determine whether the read data causes events of interest to be fired by the notification service 708 by comparing them to an acceptable range; and periodically send updates of dynamic data to the domain controller 514 via the notification service 708.
[0044] Node controller 502 may further include a command listener service 710. Command listener service 710 may receive messages or commands from domain controller 514 through a connection-oriented session that has been previously created. Command listener service 710 may queue incoming requests and fulfill them one by one at a time. Accordingly, the methods performed by command listener service 710 to perform management and monitoring operations need not be thread-safe. In one embodiment, more than one connection may be allowed in the queue.
[0045] Additionally, node controller 502 may include a serial console service 712. Serial console service 712 may operate in two modes. The first is buffered mode, in which node controller 502 collects data from the console port of server 107 and stores it in a rotating buffer. The second mode is interactive mode, in which end user 518 interacts with the serial console of server 107 via node controller 502. The interactive mode implementation emulates end user 518 as if they were directly connected to a serial port of serial console service 702, although in reality, any communication between end user 518 and serial console service 512 must go through domain controller 514 and node controller 502.In one embodiment, the buffered mode of operation may be the default mode of operation of the serial console service 702. The buffer may have a FIFO design, dropping the older data bytes to allow new bytes to be added to the top of the buffer.
[0046] A log service 714 may be further provided to receive messages from other components of node controller 502 and store them in a physical medium, such as an EEPROM. Node controller 502 may further include a monitoring service 716 for each sensor that monitors a system property of interest (e.g., temperature, electrical power consumption, voltage, current, etc.). In one embodiment, monitoring service 716 may continuously query hardware abstraction layer 704 for data for each managed hardware 718. Monitoring service 716 may retain the last value read from the sensor and may fire events to notification service 708.For example, if the temperature sensor (not shown) indicates a temperature that exceeds a predetermined safety threshold, the monitoring service 716 may fire an event to the notification service 708 informing the notification service 708 of this fact. In one embodiment, potential system errors may be reduced by having the monitoring service 716 store a value of a quantity of interest that is an average of a number of sensor readings over a predetermined time interval. In one embodiment, the sensor data may be compared with an "acceptable range" for a particular sensor to determine if a threshold has been reached. The monitoring service 716 may forward the sensor data to the domain controller 514 and / or other receivers (e.g., the fan controller 508) at a predetermined frequency.In one embodiment, the monitoring service 716 may further interact with the BMC of the sleds 106 to collect data and / or to push data.
[0047] As those skilled in the art will appreciate, with the benefits of the present invention, node controller 502 may manage a number of hardware components 718, including, but not limited to, a motherboard 718a, physical location bus / ports 718b, LEDs 718c, sensors 718d, and EEPROM / Flash 718e. However, to create flexible and maintainable code, node controller services 502 may not directly interact with the system hardware to be managed. Instead, node controller services 502 may utilize a hardware abstraction layer 704 that abstracts the hardware. For example, if command listener service 710 receives a command to turn off an LED 718c, command listener service 710 may send a request to hardware abstraction layer 704. The hardware abstraction layer 704 knows the physical medium and protocol for managing the LED.As a result, if the hardware running node controller 512 changes, only hardware abstraction layer 704 and perhaps network abstraction layer 702 need to be changed, while the other system components remain essentially the same. Node controller 502 is much more cost-effective than a full-featured baseboard management controller, yet it provides the most critical capabilities that hyperscale data center customers may desire.
[0048] Now returning to the Fig. 5 can the I 2 C signals (I2C signals) 504 are used to indicate to each carriage 106 its position in the rack 104. In addition, the I 2C signals 504 can be used as a backdoor for sending or receiving data in the event that the node controller 502 cannot obtain an IP address or the Ethernet switch 506 is damaged. In the backplane of each rack 104, there may be a switch 506 that creates an Ethernet network that allows the node controller 502 to communicate with other devices.
[0049] As in the Fig. 5, the system may include one or more power modules 510 that supply electrical power to one or more chassis racks 104. The power module 510 may include a power distribution unit (“PDU”) that receives electrical power from the data center and / or the AC receptacles that power the third-party components and provide electrical power to the chassis rack 104. The power module 510 may be connected to a power module controller 512 that provides management and monitoring capabilities for the power module 510, which provides a PDU, PSUs, and AC receptacles, etc. The power module controller 512 may communicate with a domain controller 514 via the management network. Accordingly, the system may include a shared power subsystem that distributes electrical power from the power module 510 to one or more chassis racks 104.The operating modes of the shared energy system are described in more detail in connection with the . Fig. 9 discussed.
[0050] In an exemplary embodiment, the chassis 104 may include a battery backup 116. The battery backup 116 may provide DC electrical power to the servers 107 in the event of a PDU failure. The power module controller 512 provides management and monitoring of the battery backup 116. The power module controller 512 may extract only the most critical settings and metrics provided by the battery backup 116 (e.g., battery status, remaining time, etc.) and display them to the end user 518. Alarms and / or events generated by the battery backup 116 may also be forwarded by the power module controller 512.
[0051] As will be discussed in detail below, the domain controller 514 may be operable to perform one or more of the following functions depending on the system requirements: Displaying an inventory of all devices of the rack 104; Enabling the setting and display of rack information, such as the rack name, the type of rack, and the height of the rack (e.g., 42U), to enable inventory management; Managing the electrical power of the servers in one or more sleds 106; Monitoring the power consumed by each device in the rack, as well as the total power consumption; Monitoring the temperature of the various sensors in the controllers of the rack 104; Monitoring fan speeds;Providing an aggregation of critical measures such as maximum temperature, average temperature, device errors, etc.; Detecting the failure of any controller in the rack 104 and other critical conditions; enabling the controllers in the rack 104 to be updated without negatively impacting system performance; Maintaining a history of sensor data in a database and providing statistical performance data; Enable capping of electrical power at the rack level when the total power consumption of the rack exceeds a predetermined threshold, when there is a power supply failure, or to adjust system workloads.
[0052] The domain controller 514 may be connected to a switch 516, which is used to aggregate the switches in the rack 104. An end user 518 may manage any device in the rack using the domain controller 514. This includes power management and monitoring, sensor monitoring, serial over LAN, detection of critical alarms in the rack, and / or other system features desired to be monitored or controlled.
[0053] The operation of the domain controller 514 is described in more detail with reference to the Fig. 8. As described in the Fig. As shown in Figure 8, the two main components of the domain controller 514 are the managers 802 and the interfaces 804. A manager is a module responsible for managing and monitoring a specific part of the system, and interfaces are the code that provides management and monitoring capabilities to the end users 518. The managers 802 may contain objects stored in a database. There may be an object corresponding to each device in the system, from the rack 102 to the domain controller 514 itself. Objects have properties and methods. For example, a rack object may have properties such as a maximum temperature, maximum power consumption, etc., and methods such as power management (on / off). The objects may be stored in tables of a database that contains the most current data.
[0054] The interfaces 804 receive commands from the end user 518 and communicate with the appropriate manager 802 to satisfy the requests. Accordingly, the interfaces 804 and the managers 802 are separated so that, for example, code that reads a performance measurement from a node controller 502 has nothing to do with the code that enables the domain controller 514 to reboot.
[0055] The managers 802 may include a device manager 806. The device manager 806 may be communicatively coupled to a cache 808 for sensor data provided by the low-level controllers. A single domain controller 514 may interact with many low-level controllers. For example, the device manager 806 may receive sensor data from the sleds 506, the fans 108, the chassis power supply 114, and the battery backup 116. The low-level controllers may push data to the device manager 806 of the shelf controller 514. The device manager 806 may store this data in a cache 808 so that it can be quickly retrieved when the end user 518 requests monitoring data.Additionally, the device manager 806 may store data in a database that allows the user 108 to dump historical data and allows the device manager 806 to provide the user 518 with statistical data regarding system performance. For example, in an exemplary embodiment, sensor data from each subordinate controller may be collected in a central cache 808. After a predetermined sampling interval, the entire cache contents 808 may be output to the database. Additionally, the cache may provide current monitoring data required by a consumer. For example, a request by the end user 518 regarding the real-time power consumption of a sled 106 may be satisfied by the cache 808 without requiring the device manager 806 to send a TCP command to the node controller 502.
[0056] In the unlikely event that domain controller 514 receives a packet from a subordinate controller that has not logged on, domain controller 514 will generate an event, examine the underlying User Datagram Protocol packet, obtain an IP address for the subordinate controller, and send a command to get controller information so that cache 808 can be updated. As those skilled in the art will appreciate, with the benefits of this disclosure, this should only occur if a subordinate controller has logged on to domain controller 514 and the domain controller goes offline before the update is sent to the second redundant domain controller 524.
[0057] The subordinate controllers (e.g., node controllers 502) have the capability of executing only one command at a time. In contrast, for scalability reasons, more than one command may be executed concurrently by a domain controller 514. In one embodiment, a component of the device manager 806 of the domain controller 514 may include a task pool architecture such as that used by web servers available from the Apache Software Foundation, located in Delaware, to enable the execution of more than one command at a time. Specifically, when using a task pool architecture, a set of threads may operate in parallel to execute a set of commands. For example, the electrical power of 100 nodes may be managed by having 10 threads manage the electrical power of 10 nodes.
[0058] In an exemplary embodiment, if cache 808 detects that a subordinate controller has not updated its data in a timely manner, it may send a "getsensordata" signal to the specific subordinate controller. The amount of time that may elapse before a "getsensordata" signal is sent to the specific subordinate controller may be preset by user 518 depending on system requirements. If delivery of the "getsensordata" signal to the particular subordinate controller fails, or if cache 808 does not receive a responsive signal from the subordinate controller, cache 808 may remove the old data related to that subordinate controller and generate an event to provide notification of the problem.
[0059] In an exemplary embodiment, the domain controller 514 may further include a notification manager 810. The notification manager 810 acts as a "container" for the events and alarms in the system, which are queued 811 and delivered to the notification manager 810. For example, the notification manager 810 may contain information that "the system has been started" or that "the temperature sensor in node 1 exceeds a critical threshold." The notification manager 810 is responsible for sending events of interest (e.g., a temperature above the threshold, a system has been initiated, etc.) to various destinations.In one embodiment, the notification manager 810 may intercept the events and / or alarms from a Simple Network Management Protocol (SNMP) trap 812, which may be used to monitor network-connected devices for conditions requiring administrative attention. The operation of an SNMP trap is well known to those skilled in the art and, therefore, will not be discussed in detail herein. Similarly, those skilled in the art will appreciate that, with the benefits of this disclosure, the notification manager 810 may distribute events and / or alarms to other destinations, such as, for example, a log, a system log (syslog) 814, or other suitable destinations.The SNMP trap 812 and / or the system log 814 may be used to notify an end user 514 of the events and / or alarms contained in the notification manager 810 through a user interface 816 and / or other distributors 818.
[0060] In one embodiment, the domain controller 514 may further include a security manager 820. The security manager 820 is responsible for authentication and / or role-based authorization. Authorization may be performed using a local or remote data directory. In one embodiment, the local data directory may operate under the Lightweight Directory Access Protocol ("LDAP"). The data directory may contain information about users on a local LDAP server 822 and may be extended to add additional information if / when needed. By default, the system may include a local LDAP server 822 with local users (e.g., an administrator). However, an end user 518 may add another LDAP server or similar customer directory server 823 so that the domain controller 514 understands additional users.Accordingly, in an exemplary embodiment, the domain controller 514 may have three users by default: a guest, an administrator, and an operator. The information may be stored in the local LDAP server 822. However, an end user 518 may have its own customer directory server 823 along with hundreds of users. The end user 518 should be able to connect its own customer directory server 823 to the domain controller 514 so that the domain controller 514 can then be used by any of the hundreds of users. The information for most users may be stored in a local LDAP data directory (e.g., OpenLDAP).As those skilled in the art will appreciate, with the benefits of this disclosure, if a domain controller 514 is running under the Linux system, the Linux system must be aware that the user information is stored in the local LDAP data directory and must enable Secure Shell (SSH) or Telnet authorization for the users via LDAP.
[0061] Each system manager must check the security manager 820 to determine whether an action can be performed, thereby enabling role-based access control. In one embodiment, the system may allow only two roles: (1) a guest role with read-only permissions and (2) an administrative role with read / write permissions.
[0062] Additionally, the security manager can set the firewall and restrict traffic going out of and into the domain controller 514. In one embodiment, the operations of the system can be simplified by having the security manager 820, which allows all outgoing traffic while restricting incoming traffic.
[0063] In one embodiment, the domain controller 514 may include a domain controller manager 824 that is responsible for managing the domain controller 514 itself. The functions of the domain controller manager 824 may include, for example, networking the domain controller 514, rebooting the domain controller 514, etc. Additionally, the domain controller manager 824 may allow querying log entries from the underlying file system.
[0064] The domain controller 514 may further include a redundancy manager 826. The redundancy manager 826 is responsible for sending and / or receiving "heartbeats" from the domain controllers in the network, such as the secondary domain controller 524. It is the job of the redundancy manager 826 to ensure that if one domain controller dies, another takes over without interruptions to system performance.
[0065] In one embodiment, domain controller 514 may be operable to act as a Trivial File Transfer Protocol ("TFTP") for data transfers, such as when file updates are made. Similarly, domain controller 514 may be operable to act as a Dynamic Host Configuration Protocol ("DHCP") for dynamic IP address configuration when the controller is unable to acquire a physical location. Additionally, domain controller 514 may be operable to act as a Simple Network Time Protocol ("SNTP") server to synchronize the time for all controllers on the network.
[0066] In addition to the managers 802, the domain controller 514 includes interfaces 804. In one embodiment, the domain controller 514 may include a scriptable Command Line Interface ("CLI") 828. In one embodiment, the Command Line Interface may be written with similar features to the Systems Management Architecture for Server Hardware ("SMASH") / Communication Link Protocol ("CLP"). The overall system capabilities may be represented by the scriptable CLI 828. The scriptable CLI 828 may communicate with an end user 518 using the SSH or Telnet protocol.
[0067] With the serial console service 712 in buffered mode, the end user 518 can log into the domain controller 514 to access the CLI 828. In the CLI 828, the end user 518 can type the command to access the buffered data. In response, the CLI 828 performs a task in the device manager 806. The device manager 806 can then send a TCP / IP message to the appropriate node controller 502 requesting buffered serial data. The node controller 502 will then generate a reply message and place its FIFO buffered data in this reply. The message will be received by the device manager 806 over the network, and the device manager 806 will reply to the CLI 828 with the data. The data can then be displayed by the CLI 828. When in buffered mode, the transfer of serial data from the motherboard to the FIFO of the node controller 502 is never interrupted.
[0068] In one embodiment, the serial console service 712 may be operable in interactive mode, allowing an end user 518 to interact with a server 107 through its serial port. In this embodiment, the end user 518 may log into the domain controller 514 to access the CLI 828 via SSH or Telnet. The end user 518 may then type the command to start an interactive session with a server 107 in a sled 106. In doing so, the CLI 828 performs a task in the device manager 806. The device manager 806 sends a TCP message to the appropriate node controller 502 to request the start of the interactive session. The node controller 502 may then acknowledge the command and respond to the domain controller 514 that it is ready.Additionally, the node controller 502 can spawn a thread that sends data to and receives data from the Universal Asynchronous Receiver / Transmitter ("UART"). The device manager 806 responds to the CLI 828 that the connection is ready, and the CLI 828 starts a TCP connection to the node controller 502 using the port specified for sending and receiving data. Whenever a character is received, it can be forwarded to the node controller 502, which will then forward the received character to the serial port of the designated server 107. The node controller 502 can then read the serial port of the server 107 and send the response back through the TCP connection to the CLI 828. The thread / process at the device manager 806 can then post the data to the CLI 828. The end user 518 can exit the interactive session by entering appropriate commands into the CLI 828.If buffered mode is enabled, it will not interfere with the interactive session. Rather, it should behave normally and log the output of the serial console service 712. Furthermore, since the domain controller 514 has a serial port, the customer can access the CLI 828 through this port and execute any CLI commands, including serial over LAN commands to a server 107.
[0069] The interfaces 804 of the domain controller 514 may further include an SNMP 830, which may be used to perform basic system operations such as managing the electrical power of the nodes, reading inventory, etc.
[0070] An Intelligent Platform Management Interface ("IPMI") 832 may enable a user 518 to send IPMI or Data Center Manageability Interface ("DCMI") messages through a Local Area Network ("LAN") to the domain controller 514. The domain controller 514 may provide IP aliasing to expose multiple IP addresses to the network, each associated with a specific sled 106. The message is received by the domain controller 514 and forwarded to the appropriate sled 106. The node controller 502 may process the raw IPMI packet contained in the Remote Management and Control Protocol+ ("RMCP+") message, and each IPMI software stack is processed by the domain controller 514.
[0071] An MPMI interface may further exist for each rack 102 in the domain controller 514 and may provide OEM commands for rack-level management. For example, rack-level management may include listing the inventory of the chassis 104 in a rack 102, including the sleds 106 contained therein, the positions of the sleds 106 within the chassis 104, the IPMI address of the sleds 106 to be managed, and the status of the sleds 106. Additionally, rack-level management may include information from fan controllers 508, such as, for example, the status of each fan 108 and / or the speed of each fan 108.Rack-level management may further include information about the power module controllers 512, such as the status of each PDU, the electrical power consumed, as well as an indication of critical actions of the subrack 104, such as total power consumption and maximum temperature.
[0072] The domain controller 514 may further include SMASH interfaces 834. SMASH is a standard management framework that may be placed above the manager 802. As those skilled in the art will appreciate, SMASH uses an object-oriented approach to defining management and monitoring capabilities of a system and uses "providers" to get data from the management system into the object-oriented framework. One advantage of using SMASH interfaces is that they allow the use of standard user interfaces such as SMASH / CLP 836 for command line interfaces and Common Information Model ("CIM") / Extensible Markup Language ("XML") or Web Service Management ("WS-MAN") 838 for web services.
[0073] In one embodiment, an operating system watchdog 840 may continuously check the status of various system components and restart the necessary components in the event of a failure or crash.
[0074] In one embodiment, the domain controller 514 may be responsible for enforcing a power cap if one has been set by the user as part of the rack-level power capping policy. In this embodiment, the power monitoring sensors (not shown) may be updated at a predetermined frequency. A power threshold may perform exception control actions, as required by periodically cycling a power option or by a log, if power consumption exceeds the threshold limit for a specified period of time. The exception time limit may be a multiple of the power monitoring sample time. During operation, the user 518 may define a predetermined power cap for the rack 104. The domain controller 514 may then send a message to the node controller 502 to initiate the power cap.This message may include a power threshold, an exception time limit, the action to be taken if the exception time limit is exceeded, and an emergency time limit. The system may then be set to cap the threshold or simply log the event of the threshold being exceeded. As those skilled in the art will appreciate, with the benefits of the present disclosure, the threshold may be referred to as an average power consumption over a predetermined period of time. In one embodiment, if the power consumption exceeds the threshold, a notification may be sent to the domain controller 514. If the power consumption falls below the threshold before the time limit expires, the node controller 502 will take no further action.However, if the time limit expires, the node controller 502 may enforce a cap or trigger a notification depending on the instructions it received from the domain controller 514. If the capping procedure implemented by the node controller 502 is successful, the system may continue operating. However, if the emergency time limit is reached and power consumption has not decreased below the threshold, the servers 107 are shut down. In one embodiment, the node controller 502 may store the power cap settings in flash memory so that the settings are retained even after a reset. The domain controller 514 may then enable or disable the system's power capping capabilities.Accordingly, an end user 518 can enable or disable the energy capping and / or assign the various energy capping parameters through the CLI 828 and the domain controller 514.
[0075] In one exemplary embodiment, rack-level blind capping may be used to allocate the power cap to the servers 107. In this embodiment, the capping is divided evenly among all servers 107 in the rack 102. This method is useful when all servers have similar characteristics and provide similar functionality. In another embodiment, rack-level fair capping may be used to allocate the power cap to the servers 107. In this embodiment, the power capping is enforced by allowing the redistribution of power among the servers 107 to avoid as much power capping as possible for the servers that are more heavily utilized (generally, those that consume more power).This is a continuous process, and it's a good approach to prevent performance degradation of the most critical servers, although the performance of the least power-consuming servers will be most affected. As those skilled in the art will appreciate, with the benefits of this disclosure, in any method, when a server can no longer be capped (i.e., further attempts to reduce power consumption would fail), it should be shut down sooner, thus guaranteeing the power budget.
[0076] In examples where the end user 518 has specific performance goals for an application (e.g., request response time), power capping may be used to reduce power in the servers 107 while maintaining the performance goals, ultimately reducing the operating expenses of the rack 104. Accordingly, the end user 518 may first query the power consumption in a rack 102 or set of servers 107 and cap the system to reduce power consumption. The end user 518 may then measure the application's performance under the specified power scheme. This process may be performed until an optimal performance and power capping configuration is identified.
[0077] In one exemplary embodiment, the end user 518 may apply a cap to a server 107. In another embodiment, a group-level blind cap may be used to determine the power cap for the system components. In this embodiment, once the optimal power cap has been identified through experimentation, the same cap may be applied to one or more servers in the rack 102 (it is expected that the servers are running the same applications used to determine the optimal cap). Because the scriptable CLI 828 allows an end user 518 to set power caps at a server level and read the power consumption of various devices in a rack 102, the end user could control the power capping process from an external server.
[0078] In some examples, it may be desirable to apply energy capping in the event of critical failures in the cooling system. For example, if the inlet temperature rises drastically, throttling the system components via capping could help temporarily reduce the system temperature without the need to wait for the internal thermal trip. Specifically, the end user 518 can preset a desired percentage reduction in energy consumption in the event that the thermal sensor reading exceeds a certain temperature. Energy consumption can then be reduced accordingly in the event of a thermal emergency.
[0079] In one embodiment, the end user may estimate the power consumption of a server 107 and / or the total power consumption in a rack 102, including the servers, fans, switches, etc. In this embodiment, the domain controller 514 has access to current sensor information from each controller, including the node controller 502 (for server-level measurements) and the power module controller 512 (for PDU measurements). Accordingly, the total power consumed by a rack 102, a chassis 104, or a server 107 at a given time may be calculated. Additionally, the end user may use the scriptable CLI 828 to read the power consumption of individual servers 107 and use these readings to perform calculations on an external server.
[0080] Now the Fig. 9, a shared power system in accordance with an exemplary embodiment of the present invention is generally designated by the reference numeral 900. In this exemplary embodiment, there are five sleds 106 per shelf 104. However, as will be appreciated by those skilled in the art, with the benefits of this disclosure, the methods and systems disclosed herein may be applied to a varying number of sleds 106 per shelf 104. In one embodiment, the shared power system 900 enables the sharing of 12 V from a PDU 902 to one or more shelfs 104. As shown in the Fig. 2 and Fig. 3, the sleds 106 may be connected to a common backplane 112 in the rack 104 for power and management purposes. In one embodiment, two or more 4U racks 104 may share a PDU 902, including 1+N PSUs 904, thereby creating a power domain. Each rack 102 may include two or more power domains. Further, third-party products, such as switches, may be provided in front of the PDU 902.
[0081] In one embodiment, one bus bar 906 may be provided per chassis 104 for distributing electrical power. Specifically, since the chassis 104 are located above or above other chassis 104 or the power module 510, the bus bars 906 in the rear wall of the chassis 104 may be used to distribute electrical power. In another exemplary embodiment, cables may provide directly connected distribution of electrical power from the power module 510 to each chassis 104.
[0082] The Fig. 10 illustrates the connection of a subrack 104 to a PDU 902. As those skilled in the art will appreciate, with the benefits of this disclosure, more than one subrack 904 may be connected to a PDU 902. Further, a PDU 902 may include one or more PSUs 904. In one embodiment, the PDU 902 may include N+1 PSUs 904 for redundancy purposes, where N is the number of sleds 106 powered by the power supply. Specifically, using N+1 PSUs 904 allows the PDU 902 to satisfy the load requirements while providing redundancy such that if one PSU 904 dies, an additional PSU 904 remains to take over the load. As shown in the Fig. As shown in Figure 10, the bus bar 906 of the PDU 902 may be connected to the bus bar 908 of the subrack 104 by one or more power cables. The bus bar 908 may then supply electrical power to the sled 106 through the backplane 112.
[0083] Although example embodiments are described in connection with servers in a rack, those skilled in the art will appreciate that the benefits of this disclosure of the present invention are not limited to servers and may be used in connection with other information handling systems, such as data storage devices. Furthermore, the system and methods disclosed herein are not limited to systems including one rack and may be used in connection with two or more racks. As those skilled in the art will appreciate, in multi-rack systems, the benefits of this disclosure enable the domain controller 514 to be scalable and may support multi-rack management capabilities by connecting management switches of other racks to the aggregation switch 516.
[0084] Although the present disclosure has been described in detail, it should be understood that various changes, substitutions and alterations may be made thereto without departing from the spirit and scope of the invention as defined in the appended claims.
Claims
[1] A modular information processing system framework comprising: a rack containing at least one subrack; at least two carriages placed in the subrack; wherein each of the carriages comprises at least one information processing system; a fan placed in the chassis to cool the information processing system; a fan controller that is communicatively connected to the fan; wherein the fan controller manages the operation of the fan; at least two node controllers, each of the at least two node controllers being associated with a different one of the at least two carriages; wherein each of the node controllers manages the operation of the associated sled; a power module for supplying electrical power to the information processing system; a power module controller for managing the operation of the power module; a primary domain controller communicatively coupled to at least one of the fan controller, the node controllers, and the power module; wherein the primary domain controller manages the operation of at least one of the fan controller, the node controllers, and the power module; and a pooling switch in each of the racks connected to a central management switch at the primary domain controller, the pooling switch providing a redundant connection point for a user through an interface. [2] The system of claim 1, wherein the primary domain controller provides a user interface for the modular information handling system framework. [3] The system of claim 2, wherein the primary domain controller is operable to display information relating to the performance of the modular information handling system framework to the user. [4] The system of claim 2, wherein the primary domain controller enables a user to control the performance parameters of the modular information handling system framework. [5] The system of claim 1, further comprising a secondary domain controller, wherein the secondary domain controller manages the operation of at least one of the fan controller, the node controller, and the power module when the primary domain controller becomes inoperative. [6] The system of claim 5, further comprising a management device, wherein the management device comprises one of the primary domain controller, the secondary domain controller, and a management switch. [7] The system of claim 1, wherein the domain controller is communicatively connected to at least one of the power module controller, the fan controller, and the node controller through a management network. [8] The system of claim 7, wherein the management network is an Ethernet network. [9] The system of claim 1, wherein the information handling system is selected from the group comprising a server and a data storage device. [10] The system of claim 1, further comprising one or more sensors for monitoring operating conditions of at least one of the sled, the fan, and the power module. [11] The system of claim 10, wherein the one or more sensors are selected from a group comprising a temperature sensor and an energy monitoring sensor. [12] The system of claim 1, wherein the primary domain controller comprises one or more managers and one or more interfaces. [13] The system of claim 10, wherein the primary domain controller comprises: a device manager that is communicatively connected to one or more sensors; wherein the device manager receives sensor data from the one or more sensors; a domain controller manager to manage the primary domain controller; a security manager to authenticate connections to the primary domain controller; and a notification manager for monitoring sensor data received by the device manager. [14] The system of claim 13, wherein the notification manager generates a notification when sensor data indicates the occurrence of an event of interest. [15] The system of claim 14, wherein the event of interest is selected from the group comprising sensor data indicating a temperature exceeding a threshold temperature value and sensor data indicating energy consumption exceeding a threshold value. [16] The system of claim 14, wherein the notification manager generates the notification for an end user through a user interface. [17] The system of claim 5, wherein the primary domain controller comprises a redundancy manager, the redundancy manager enabling communications between the primary domain controller and the secondary domain controller. [18] The system of claim 17, wherein the operation of the primary domain controller is transferred to the secondary domain controller if the secondary domain controller is unable to communicate with the redundancy manager. [19] The system of claim 13, wherein the device manager includes a cache for buffering sensor data received from the one or more sensors. [20] The system of claim 19, wherein the sensor data is moved from the buffer to a permanent memory at a predetermined frequency. [21] The system of claim 1, wherein the fan controller comprises: an identification service for identifying the physical location of the fan controller in the system; a notification service operable to send messages from the fan controller to the primary domain controller; a command listening service for receiving messages from the primary domain controller; a monitoring service; wherein the monitoring service tracks data from one or more sensors associated with the fan; a dynamic fan control for regulating the speed of the fan based on information available from the monitoring service; a log service operable to receive and store messages from components of the fan controller; and a heartbeat signal to determine whether the fan controller is operating. [22] The system of claim 21, wherein the monitoring service generates a signal to the notification service when data from the one or more sensors indicate an event of interest. [23] The system of claim 21, wherein the dynamic fan control comprises a proportional-integral-derivative controller. [24] The system of claim 21, wherein the heartbeat signal instructs the fan to operate at maximum speed when the fan controller is inoperative. [25] The system of claim 1, wherein the node controller manages electrical power for the sled. [26] The system of claim 1, wherein the node controller comprises: an identification service for identifying the physical position of the node controller in the system; a notification service operable to send messages from the node controller to at least one of the primary domain controller and the fan controller; a command listener for receiving messages from the primary domain controller; a serial console service for connecting to the information processing system; a monitoring service; wherein the monitoring service tracks data from one or more sensors associated with the sled; and a log service operable to receive and store messages from one or more components of the node controller. [27] The system of claim 26, wherein the monitoring service is operable to generate a signal to the notification service when data from one or more sensors indicate an event of interest. [28] The system of claim 1, wherein the node controller is communicatively coupled to the fan controller. [29] The system of claim 1, wherein at least one of the fan controller and the node controller includes a configuration file containing its operating parameters. [30] The system of claim 1, wherein the primary domain controller is operable to configure the fan controller configuration file and the node controller configuration file. [31] A modular rack system comprising: a plurality of subracks placed in one or more racks; a plurality of carriages placed in each subrack; each carriage comprising an information processing system; at least two node controllers, each node controller associated with a different one of the plurality of sleds, and each node controller managing the operation of an associated sled; a federation switch in each of the racks connected to a central management switch at the primary domain controller, the federation switch providing a redundant connection point for a user through an interface; a shared fan module for cooling the plurality of sleds in each subrack; a shared power module for supplying electrical power to the one or more carriages in one or more chassis; and a shared management module for managing the operation of the majority of the racks. [32] The system of claim 31, wherein the shared fan module comprises one or more fans and a fan controller for controlling operation of the shared fan module. [33] The system of claim 32, wherein the fan controller comprises: an identification service for identifying the physical position of the fan controller within the system; a notification service operable to send messages from the fan controller to the primary domain controller; a command listening service for receiving messages from the primary domain controller; a monitoring service; wherein the monitoring service tracks data from one or more sensors associated with the fan; a dynamic fan control for regulating the speed of the fan based on information available from the monitoring service; a log service operable to receive and store messages from components of the fan controller; and a heartbeat signal to determine whether the fan controller is operating. [34] The system of claim 33, wherein the monitoring service is operable to generate a signal to the notification service when data from the one or more sensors indicates an event of interest. [35] The system of claim 31, wherein the shared power module comprises a power distribution unit and a power module controller for controlling operation of the shared power module. [36] The system of claim 35, wherein the power supply unit comprises one or more power supply units. [37] The system of claim 31, wherein the shared management module comprises a primary domain controller. [38] The system of claim 37, wherein the primary domain controller enforces a power capping policy. [39] The system of claim 37, wherein the primary domain controller tracks the power consumption of one or more components of the modular rack system. [40] The system of claim 37, further comprising a secondary domain controller that replaces the primary domain controller when the primary domain controller becomes inoperative. [41] The system of claim 37, wherein the primary domain controller comprises: a device manager that is communicatively connected to one or more sensors; wherein the one or more sensors monitor the operating conditions of at least one of a sled, a component of the shared fan module, and a component of the shared power module; wherein the device manager receives sensor data from the one or more sensors; a domain controller manager that is operational to manage the primary domain controller; a security manager operable to authenticate connections to the primary domain controller; and a notification manager operable to monitor sensor data received by the device manager. [42] The system of claim 41, wherein the notification manager is operable to generate a notification when sensor data indicates the occurrence of an event of interest. [43] The system of claim 37, further comprising a node controller associated with a sled. [44] The system of claim 43, wherein the node controller comprises: an identification service for identifying the physical position of the node controller within the system; a notification service operable to send data from the node controller to at least one member of the primary domain controller group and the shared fan controller; a command listener for receiving messages from the primary domain controller; a serial console service for connecting to the information processing system in the sled; a monitoring service; wherein the monitoring service tracks data from one or more sensors associated with the sled; wherein the monitoring service is operable to generate a signal to the notification service when data from the one or more sensors indicate an event of interest; and a log service that is operational to receive and store messages from one or more components of the node controller. [45] The system of claim 31, wherein the shared fan module, the shared power module, and the shared management module are communicatively linked via a management network. [46] The system of claim 45, wherein the management network is an Ethernet network. [47] The system of claim 31, wherein power is distributed to the carriages through a chassis backplane.
Citation Information
Patent Citations
Blade server system
US20060007651A1
Enclosure for computer peripheral devices
US20060041783A1
Modular node assembly for rack mounted multiprocessor computer
US6025989A