Distributed BMC Architecture for Server Board Redundancy
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Server boards with a single centralized baseboard management controller (BMC) are costly and prone to failure, leading to entire board failure, as they require dedicated memory, communication ports, and power supply, and lack redundancy for continuous operation.
Innovation Solution
Implementing a distributed BMC architecture where multiple service processors operate together, with one as a primary BMC and others as slave BMCs, monitoring each other's health and implementing remedial actions, allowing the system to continue functioning even if one BMC fails, without the need for additional monitoring devices.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If a single centralized BMC is implemented, then the server board can be monitored and controlled, but the cost increases and the risk of single-point failure increases
Solution Approach 1:
The patent divides the single centralized BMC into multiple distributed service processors (SPs), each capable of independent operation. Each SP monitors and controls specific components or devices on the server board, eliminating the single-point failure risk while distributing the monitoring and control functions across multiple independent units.
Solution Approach 2:
Each service processor is configured with local monitoring and control capabilities tailored to its specific responsibilities. Each SP has dedicated memory, communication ports, and power supply circuitry optimized for its local functions, allowing independent operation and reducing the impact of failures on the entire system.
2Reliability
If dedicated memory, communication ports, and power supply circuitry are allocated to BMC, then BMC functionality is ensured, but the server board cost increases
Solution Approach 1:
The service processors are designed as multi-functional units that can perform both BMC monitoring/control functions and general-purpose computing tasks. By making the SPs universal, the dedicated hardware resources serve dual purposes, reducing the overall hardware footprint and cost while maintaining reliable BMC functionality.
Solution Approach 2:
The patent combines BMC functionality with general-purpose service processor functions into a single integrated unit. Each SP integrates monitoring, control, communication, and power management capabilities along with general computing functions, eliminating the need for separate dedicated BMC hardware and reducing overall component count and cost.
3Ease of operation
If a single BMC monitors all components, then centralized control is achieved, but the system lacks redundancy for continuous operation
Solution Approach 1:
The patent implements a dynamic distributed architecture where service processors can assume different roles (primary or secondary) based on operational conditions. Each SP continuously monitors the status of other SPs and can dynamically take over monitoring and control functions if another SP fails, maintaining centralized control while providing redundancy for continuous operation.
Solution Approach 2:
Each service processor continuously monitors the health and operational status of other SPs through inter-processor communication. This feedback mechanism allows the system to detect failures and automatically redistribute monitoring and control responsibilities, ensuring continuous operation while maintaining coordinated control across all components.
Data Source
AI summary
A server board includes first and second devices. A first service processor of the first device operates as a master baseboard management controller of the server board, and monitors a communication channel for alive messages from a plurality service processors. A second service processor operates as a secondary baseboard management controller, and sets a second timer to a first value. In response to a determination that the second timer has expired based on a first value: the second service processor to start a switchover process, and to set the second timer to a second value based on an alive message period. In response to a primary alive message not being received from the first service processor prior to the second timer expiring based on the second value, the second service processor to reset first service processor and to operate as the master baseboard management controller.


