Blade Server Management Module Failover via Impeachment Consensus

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current blade server chassis systems have limited failover capabilities, as the standby Management Module (MM) cannot determine when to take over from a failed or overloaded primary MM, especially when the primary MM is too busy or not properly servicing interrupts.

Innovation Solution

A computer-implemented method where each server blade evaluates the performance of the primary MM, and if a threshold number of blades determine it is not meeting minimum standards, the secondary MM takes over, utilizing a performance-based switching logic to impeach the primary MM and manage the server blades.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If the standby MM waits for explicit failure notification from the primary MM, then the system maintains simplicity in failover triggering, but the standby MM cannot detect performance degradation when the primary MM is overloaded or unresponsive

Engineering Contradiction:
Improvefailover detection accuracyVSAvoidfailover mechanism complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent implements a feedback mechanism where server blades continuously monitor primary MM performance by evaluating whether the primary MM responds to their service requests within expected timeframes. Each blade provides feedback about its interaction with the primary MM, and when a threshold number of blades report failures to communicate with the primary MM, this triggers automatic failover to the standby MM. This resolves the contradiction by enabling reliable detection of performance degradation through distributed feedback from multiple blades.

Inventive Principle:
Principle #23Feedback

2Reliability

If the system implements automatic failover based on performance monitoring, then service continuity is improved, but the risk of premature or incorrect failover increases

Engineering Contradiction:
Improveservice continuityVSAvoidrogue failover risk
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The patent requires a threshold number of server blades to report communication failures with the primary MM before triggering failover, rather than acting on a single blade's report. This partial action approach prevents premature failover due to isolated incidents while still enabling timely failover when multiple blades experience the same issue. The threshold mechanism filters out false positives while maintaining sensitivity to genuine failures, thus improving service continuity without increasing rogue failover risk.

Inventive Principle:
Principle #16Partial or excessive action

3Loss of time

If the standby MM actively monitors primary MM performance through blade evaluations, then failover timing is optimized, but the system overhead and complexity increase

Engineering Contradiction:
Improvefailover response timeVSAvoidperformance monitoring complexity
Core Design Contradiction:
Loss of timeVSDevice complexity

Solution Approach 1:

The patent implements self-service monitoring where server blades autonomously evaluate the primary MM's performance by attempting to communicate with it and determining whether responses are received within expected timeframes. Each blade independently assesses the primary MM's responsiveness without requiring centralized monitoring infrastructure. This self-service approach optimizes failover response time by distributing the monitoring function across multiple blades while avoiding the complexity of a dedicated centralized monitoring system.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS8037364B2Forced management module failover by BMC impeachment consensus
Publication Date: 2011.10.11 LENOVO INT LTD
  • US8037364B2 patent drawing
  • US8037364B2 patent drawing
  • US8037364B2 patent drawing

AI summary

A computer-implemented method, system and computer program product for managing failover of Management Modules (MMs) in a blade chassis are presented. Each server blade in the blade chassis evaluates a performance of a primary MM. If a threshold number of server blades determine that the primary MM is not meeting pre-determined minimum performance standards, then a secondary MM impeaches the primary MM and takes over the management of the server blades.