Failure Detection Engine for CEC Group Health Monitoring

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Cluster-based high availability management for central electronics complexes (CECs) is complex and challenging due to the need for frequent health status queries that can overload hardware management consoles (HMCs) and negatively affect other operations, especially when direct communication with LPARs and VIOSes is restricted for security reasons.

Innovation Solution

A failure detection engine (FDE) periodically sends REST API probes to HMCs at predefined intervals to collect health data from VIOSes and LPARs, prioritizing VIOS health monitoring and only returning data on unhealthy LPARs to reduce bandwidth and traffic, allowing for relocation decisions to be made based on health data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If frequent health status queries are sent to HMC to monitor LPAR and VIOS health, then failure detection capability is improved, but HMC load increases and other operations are negatively affected

Engineering Contradiction:
Improvefailure detection capabilityVSAvoidHMC load
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent extracts only the necessary health status information from the HMC responses. The FDE sends probes to HMC and selectively processes only the health data relevant to LPAR and VIOS status, leaving other HMC operations unaffected. This extraction approach reduces HMC load while maintaining failure detection capability.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system performs partial monitoring by focusing only on critical health parameters rather than comprehensive system monitoring. The FDE monitors only essential LPAR and VIOS health indicators, allowing frequent checks without overwhelming the HMC with excessive comprehensive queries.

Inventive Principle:
Principle #16Partial or excessive action

2Loss of information

If health data for all LPARs is returned in response packets, then complete monitoring information is provided, but bandwidth usage increases

Engineering Contradiction:
Improvemonitoring information completenessVSAvoidbandwidth usage
Core Design Contradiction:
Loss of informationVSLoss of energy

Solution Approach 1:

The HMC extracts and returns only the specific health data for LPARs that are relevant to the monitoring requirements. Instead of returning complete data sets for all LPARs, the system selectively transmits only the necessary health status information, reducing bandwidth consumption while maintaining monitoring effectiveness.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The monitoring system applies local quality by providing detailed health information only where needed (for LPARs requiring monitoring) rather than uniformly distributing data across all LPARs. This targeted approach optimizes bandwidth usage by transmitting information selectively based on local monitoring requirements.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS10713138B2Failure detection for central electronics complex group management
Publication Date: 2020.07.14 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US10713138B2 patent drawing
  • US10713138B2 patent drawing
  • US10713138B2 patent drawing

AI summary

Examples of techniques for failure detection for central electronics complex (CEC) group management are described herein. An aspect includes issuing a first logical partition (LPAR) probe to a hardware management console (HMC) of a central electronics complex (CEC) group, wherein the CEC group comprises a plurality of LPARs. Another aspect includes receiving a first response packet from the HMC corresponding to the first LPAR probe, wherein the first response packet comprises health data corresponding to a first LPAR of the plurality of LPARs. Another aspect includes storing the health data corresponding to the first LPAR in a first health data entry corresponding to the first LPAR. Another aspect includes, for a second LPAR of the plurality of LPARs that was not included in the first response packet, updating a second health data entry corresponding to the second LPAR to indicate that the second LPAR is healthy.