Failover Anomaly Detection in High Availability Clusters
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
High-availability computing systems face anomalies during failover processes due to neglect of disaster recovery protocols, leading to potential abortive failovers and compromised business continuity.
Innovation Solution
A method that involves obtaining and monitoring parameters associated with disaster recovery protocols across clusters, detecting anomalies that impact failover, and generating alerts to enable corrective actions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If disaster recovery protocols are manually monitored and administered, then system complexity is reduced, but reliability of failover deteriorates due to human error and neglect
Solution Approach 1:
The system performs self-monitoring of disaster recovery protocol parameters through automated agents that continuously check system states, compare against expected values, and detect anomalies without human intervention. The monitoring agents autonomously identify deviations in replication status, data integrity, and resource availability, enabling the system to self-diagnose potential failover issues.
Solution Approach 2:
The system implements continuous feedback loops where monitoring agents report parameter states to a central management system, which then generates alerts when anomalies are detected. This feedback mechanism ensures that deviations from expected disaster recovery states are immediately communicated to administrators, enabling timely corrective actions before failover is impacted.
2Measurement precision
If continuous monitoring of disaster recovery parameters is implemented, then detection precision of anomalies improves, but use of energy and computational resources increases
Solution Approach 1:
The monitoring system applies partial monitoring by focusing only on critical disaster recovery parameters such as replication status, data integrity checks, and resource availability. Rather than monitoring all system parameters continuously, the system selectively monitors only those parameters directly related to failover success, reducing computational overhead while maintaining adequate detection precision for critical anomalies.
3Loss of time
If manual administration of disaster recovery protocols is used, then ease of operation is maintained, but loss of time for detecting anomalies increases
Solution Approach 1:
The system autonomously performs continuous monitoring and anomaly detection without requiring manual administration. Monitoring agents automatically check disaster recovery parameter states, compare them against expected values, and detect anomalies in real-time. This self-service approach eliminates the time delay associated with manual checking while maintaining operational simplicity through automated alerting to administrators.
Solution Approach 2:
The monitoring system operates continuously rather than periodically or on-demand, ensuring that disaster recovery parameters are constantly monitored for anomalies. This continuous monitoring enables immediate detection of issues as they arise, significantly reducing the time loss associated with anomaly detection compared to manual periodic checks, while the automated nature maintains ease of operation.
Data Source
AI summary
A method, apparatus and system for improving failover within a high-availability computer system are provided. The method includes obtaining one or more parameters associated with a disaster recovery protocol of at least one resource of any of the first cluster, second cluster and high-availability computer system. The method also includes monitoring one or more states of the parameters. The method further includes detecting, as a function of the parameters and states, one or more anomalies of any of the first cluster, second cluster and high-availability computer system, wherein the anomalies are types that impact the failover. These anomalies may include anomalies associated with the disaster-recovery protocols within the first and/or second clusters (“intra-cluster anomalies”) and/or anomalies among the first and second clusters (“inter-cluster anomalies”). The method further includes generating an alert in response to detecting one or more of the anomalies.


