Confidence Measure for Distributed System Failover Readiness
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current disaster recovery solutions lack the ability to proactively assess the readiness of high availability (HA) in distributed computing systems, specifically failing to measure the success of takeover by secondary servers during disasters and the performance impact post-disaster scenarios.
Innovation Solution
A confidence measure is computed for controllers in a distributed computing system, based on veto and non-veto events, to determine HA readiness, allowing administrators to assess the likelihood of successful failover and potential performance impacts, enabling proactive configuration for disaster recovery.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If disaster recovery solutions use redundant machines to provide data protection, then data availability is improved, but the ability to proactively measure takeover success and performance impact is lost
Solution Approach 1:
The system performs preliminary actions by computing a confidence measure before a disaster occurs. This confidence measure is calculated based on the current state of the secondary server and historical performance data, allowing administrators to assess whether the secondary server is ready to successfully take over if the primary server fails. This preliminary assessment resolves the contradiction by providing measurement capability without requiring an actual disaster event.
Solution Approach 2:
The system implements feedback mechanisms by continuously monitoring the state of redundant machines and computing confidence measures that reflect their readiness status. This feedback loop provides administrators with information about the likelihood of successful takeover, enabling them to make informed decisions about system configuration and disaster recovery preparedness. The feedback resolves the measurement precision problem by providing quantitative data about takeover success probability.
2Reliability
If disaster recovery solutions fail over to a second server during disaster, then data access is maintained, but the performance impact on the second server cannot be determined
Solution Approach 1:
The system performs preliminary performance assessment by computing the confidence measure before failover occurs. This confidence measure incorporates performance metrics and capacity analysis of the secondary server, allowing administrators to predict the performance impact that would occur during and after failover. This preliminary action resolves the contradiction by providing performance visibility without requiring an actual disaster event.
Solution Approach 2:
The system enables self-service by automatically computing confidence measures and performance impact assessments without requiring manual testing or actual disaster events. The automated analysis of server capacity, workload characteristics, and historical performance data provides administrators with actionable information about expected performance impacts, resolving the productivity measurement problem through self-assessment capabilities.
3Reliability
If multiple copies of data are stored at different datacenters, then disaster recovery capability is improved, but the ability to assess HA readiness is reduced
Solution Approach 1:
The system applies universality by using a standardized confidence measure computation approach that works across multiple datacenters and server configurations. The same methodology for assessing readiness is applied regardless of geographical location or specific hardware setup, simplifying the complexity of multi-datacenter HA readiness assessment. This universal approach resolves the contradiction by providing a consistent assessment framework across distributed systems.
Solution Approach 2:
The system uses parameter changes by computing confidence measures based on varying states of servers and systems. The confidence measure dynamically adjusts based on current performance data, capacity metrics, and system state, providing a simplified yet accurate assessment of HA readiness across multiple datacenters. This parameter-based approach resolves the complexity issue by transforming complex multi-datacenter assessment into a manageable computational process.
Data Source
AI summary
Technology is disclosed for determining high availability readiness of a distributed computing system (“system”). A confidence measure (CM) can be computed for a particular controller in the system to determine whether a takeover by the particular controller from a first controller would be successful. The CM can be a percentage value. A CM of 0% indicates that a takeover would be a failure, which results in loss of access to data managed by the first controller. A CM of 100% indicates a successful takeover with no performance impact on the system. A CM between 0% and 100% indicates a successful takeover but with a performance impact. The CM can be computed based on events occurring in the system, e.g., veto and non-veto events. The CM is computed as a function of various weights and/or indices associated with the veto events and/or non-veto events.


