Failure Detection Engine for CEC Group Health Monitoring
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Cluster-based high availability management for central electronics complexes (CECs) is complex and challenging due to the need for frequent health status queries that can overload hardware management consoles (HMCs) and negatively affect other operations, especially when direct communication with LPARs and VIOSes is restricted for security reasons.
Innovation Solution
A failure detection engine (FDE) periodically sends REST API probes to HMCs at predefined intervals to collect health data from VIOSes and LPARs, prioritizing VIOS health monitoring and only returning data on unhealthy LPARs to reduce bandwidth and traffic, allowing for relocation decisions to be made based on health data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If frequent health status queries are sent to HMC to monitor LPAR and VIOS health, then failure detection capability is improved, but HMC load increases and other operations are negatively affected
Solution Approach 1:
The patent extracts only the necessary health status information from the HMC responses. The FDE sends probes to HMC and selectively processes only the health data relevant to LPAR and VIOS status, leaving other HMC operations unaffected. This extraction approach reduces HMC load while maintaining failure detection capability.
Solution Approach 2:
The system performs partial monitoring by focusing only on critical health parameters rather than comprehensive system monitoring. The FDE monitors only essential LPAR and VIOS health indicators, allowing frequent checks without overwhelming the HMC with excessive comprehensive queries.
2Loss of information
If health data for all LPARs is returned in response packets, then complete monitoring information is provided, but bandwidth usage increases
Solution Approach 1:
The HMC extracts and returns only the specific health data for LPARs that are relevant to the monitoring requirements. Instead of returning complete data sets for all LPARs, the system selectively transmits only the necessary health status information, reducing bandwidth consumption while maintaining monitoring effectiveness.
Solution Approach 2:
The monitoring system applies local quality by providing detailed health information only where needed (for LPARs requiring monitoring) rather than uniformly distributing data across all LPARs. This targeted approach optimizes bandwidth usage by transmitting information selectively based on local monitoring requirements.
Data Source
AI summary
Examples of techniques for failure detection for central electronics complex (CEC) group management are described herein. An aspect includes issuing a first logical partition (LPAR) probe to a hardware management console (HMC) of a central electronics complex (CEC) group, wherein the CEC group comprises a plurality of LPARs. Another aspect includes receiving a first response packet from the HMC corresponding to the first LPAR probe, wherein the first response packet comprises health data corresponding to a first LPAR of the plurality of LPARs. Another aspect includes storing the health data corresponding to the first LPAR in a first health data entry corresponding to the first LPAR. Another aspect includes, for a second LPAR of the plurality of LPARs that was not included in the first response packet, updating a second health data entry corresponding to the second LPAR to indicate that the second LPAR is healthy.


