Dynamic Priority List for Cluster Failover Health Monitoring
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current clustering software and health monitoring tools are not cluster-aware, failing to proactively identify the cause of application failures and provide effective corrective measures, leading to suboptimal application performance and availability.
Innovation Solution
A method and apparatus for proactively monitoring application health data by accessing performance data and event information to determine application health and dynamically selecting a target node for failover based on dynamic priority lists, ensuring optimal resource allocation and high availability.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Speed
If static priority list is used for failover, then failover process is simple and fast, but failover may target nodes affected by the same failure cause
Solution Approach 1:
The patent transforms the static priority list into a dynamic priority list that adapts based on failure cause analysis. The system monitors application health data, identifies failure causes, and dynamically reorders the priority list to exclude nodes affected by the same failure cause, thus resolving the contradiction between fast failover and reliable failover target selection
Solution Approach 2:
The system implements feedback by continuously monitoring application health data and failure causes, then using this information to dynamically adjust the priority list. This feedback loop ensures that failover decisions are based on current system state and failure patterns, improving reliability without significantly impacting failover speed
2Reliability
If clustering software monitors application status, then application availability is tracked, but software cannot identify failure causes or provide corrective measures
Solution Approach 1:
The patent applies preliminary action by proactively analyzing application health data and identifying potential failure causes before they result in complete application failure. The system continuously examines performance metrics, event logs, and health indicators to detect early signs of problems and determine root causes, enabling preventive corrective actions
Solution Approach 2:
The system introduces an intermediary health analysis component that bridges the gap between basic status monitoring and actionable failure diagnosis. This intermediary layer processes raw health data, identifies failure patterns and causes, and translates them into meaningful corrective recommendations, thus preserving both monitoring capability and diagnostic information
3Device complexity
If static failover list is used, then system complexity is low, but system cannot adapt to different failure scenarios
Solution Approach 1:
The patent implements dynamics by making the failover priority list adaptive rather than fixed. The system dynamically reconfigures the priority list based on analyzed failure causes, automatically adjusting failover behavior to suit different failure scenarios such as network failures, hardware failures, or software errors without requiring manual intervention or increasing overall system complexity
Data Source
AI summary
A method and apparatus for proactively monitoring data center health data to achieve workload management and high availability is provided. In one embodiment, a method for processing application health data to improve application performance within a cluster includes accessing at least one of performance data or event information associated with at least one application to determine application health data and examining the application health data to identify an application of the at least one application to migrate.


