Distributed BMC Architecture for Server Board Redundancy

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Server boards with a single centralized baseboard management controller (BMC) are costly and prone to failure, leading to entire board failure, as they require dedicated memory, communication ports, and power supply, and lack redundancy for continuous operation.

Innovation Solution

Implementing a distributed BMC architecture where multiple service processors operate together, with one as a primary BMC and others as slave BMCs, monitoring each other's health and implementing remedial actions, allowing the system to continue functioning even if one BMC fails, without the need for additional monitoring devices.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If a single centralized BMC is implemented, then the server board can be monitored and controlled, but the cost increases and the risk of single-point failure increases

Engineering Contradiction:
Improvesystem reliabilityVSAvoidBMC architecture complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent divides the single centralized BMC into multiple distributed service processors (SPs), each capable of independent operation. Each SP monitors and controls specific components or devices on the server board, eliminating the single-point failure risk while distributing the monitoring and control functions across multiple independent units.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Each service processor is configured with local monitoring and control capabilities tailored to its specific responsibilities. Each SP has dedicated memory, communication ports, and power supply circuitry optimized for its local functions, allowing independent operation and reducing the impact of failures on the entire system.

Inventive Principle:
Principle #3Local quality

2Reliability

If dedicated memory, communication ports, and power supply circuitry are allocated to BMC, then BMC functionality is ensured, but the server board cost increases

Engineering Contradiction:
ImproveBMC functionality reliabilityVSAvoidhardware resources
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The service processors are designed as multi-functional units that can perform both BMC monitoring/control functions and general-purpose computing tasks. By making the SPs universal, the dedicated hardware resources serve dual purposes, reducing the overall hardware footprint and cost while maintaining reliable BMC functionality.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent combines BMC functionality with general-purpose service processor functions into a single integrated unit. Each SP integrates monitoring, control, communication, and power management capabilities along with general computing functions, eliminating the need for separate dedicated BMC hardware and reducing overall component count and cost.

Inventive Principle:
Principle #5Merging (Combining)

3Ease of operation

If a single BMC monitors all components, then centralized control is achieved, but the system lacks redundancy for continuous operation

Engineering Contradiction:
Improvecentralized controlVSAvoidcontinuous operation capability
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The patent implements a dynamic distributed architecture where service processors can assume different roles (primary or secondary) based on operational conditions. Each SP continuously monitors the status of other SPs and can dynamically take over monitoring and control functions if another SP fails, maintaining centralized control while providing redundancy for continuous operation.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

Each service processor continuously monitors the health and operational status of other SPs through inter-processor communication. This feedback mechanism allows the system to detect failures and automatically redistribute monitoring and control responsibilities, ensuring continuous operation while maintaining coordinated control across all components.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS10013319B2Distributed baseboard management controller for multiple devices on server boards
Publication Date: 2018.07.03 NXP USA INC
  • US10013319B2 patent drawing
  • US10013319B2 patent drawing
  • US10013319B2 patent drawing

AI summary

A server board includes first and second devices. A first service processor of the first device operates as a master baseboard management controller of the server board, and monitors a communication channel for alive messages from a plurality service processors. A second service processor operates as a secondary baseboard management controller, and sets a second timer to a first value. In response to a determination that the second timer has expired based on a first value: the second service processor to start a switchover process, and to set the second timer to a second value based on an alive message period. In response to a primary alive message not being received from the first service processor prior to the second timer expiring based on the second value, the second service processor to reset first service processor and to operate as the master baseboard management controller.