Distributed BMC Management Plane for Datacenter Reliability
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Datacenter administrators face challenges in effective server provisioning, monitoring, and management due to centralized 'north-bound' solutions that can lead to single points of failure and require additional hardware or software expenditures, which cloud and datacenter scale users are reluctant to incur.
Innovation Solution
A distributed system using the 'Hello BMC Protocol' that interconnects baseboard management controllers (BMCs) in a datacenter, enabling a master-slave relationship and hierarchical tree topology for decentralized management, allowing each BMC to discover and configure itself within a network, proliferate seeded information, and maintain live child matrices using heartbeats.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If centralized north-bound management solutions are implemented, then server provisioning and monitoring can be managed from a single location, but a single point of failure is created and additional hardware/software costs are incurred
Solution Approach 1:
The patent segments the centralized management function into distributed BMCs at each server. Each BMC independently manages its local server, eliminating the single point of failure while maintaining manageable operation through standardized protocols. The management function is divided across multiple nodes rather than concentrated at one location.
Solution Approach 2:
Each BMC autonomously performs provisioning, monitoring, and management tasks for its local server without requiring constant intervention from a centralized administrator. The system enables self-service operations where the management plane automatically discovers servers, establishes relationships, and executes management functions independently.
2Ease of operation
If centralized north-bound management solutions are implemented, then comprehensive control is achieved, but additional hardware and software expenditures are required
Solution Approach 1:
The patent extracts the management function from separate centralized hardware/software systems and integrates it directly into the server BMC. This eliminates the need for additional UCS Managers, Fabric Interconnect switches, or centralized management software while maintaining comprehensive control capabilities.
Solution Approach 2:
The BMC serves multiple functions: it manages local server operations, participates in the distributed management plane, discovers other servers, and executes management protocols. This multi-functionality eliminates the need for separate dedicated management hardware and software systems.
3Reliability
If distributed BMC networking is implemented, then single point of failure risk is reduced, but automatic engagement and configuration protocols become more complex
Solution Approach 1:
The patent implements preliminary actions by pre-configuring BMCs with discovery protocols and management templates before deployment. The Hello BMC Protocol is pre-established, allowing servers to automatically engage and configure themselves upon power-up without requiring complex real-time negotiation or manual configuration.
Solution Approach 2:
The system uses feedback mechanisms where BMCs exchange status information, capability advertisements, and configuration data through standardized protocols. This feedback enables automatic engagement and configuration while keeping the protocol manageable through clear, structured communication cycles.
Data Source
AI summary
Presented herein are methodologies for managing servers in a datacenter or cloud service environment. A method includes sending, from a first baseboard management controller (BMC), to other BMCs in a network, a first message indicating a desire to establish a master-slave relationship; receiving a response from a second BMC from among the other BMCs, the response indicating an ability to function as a master in the master-slave relationship; sending, from the first BMC, a second message to the second BMC confirming establishment of the master-slave relationship between the second BMC and the first BMC; sending, from the first BMC to the second BMC, a request for configuration information; and in response to the request for configuration information, receiving, from the second BMC, configuration profile data representative of a configuration of a server, which is controlled by the second BMC.


