Cloud-Mediated Node State Verification for Failover Accuracy
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In cloud computing environments, computing nodes may erroneously perform switchover operations due to communication link failures or infrastructure issues, leading to incorrect assumptions about partner node failures and resulting in loss of access to storage services.
Innovation Solution
Computing nodes are configured to share operational state information through a cloud environment node state provider and/or cloud persistent storage accessible via a cloud storage service, allowing them to reliably determine whether a partner node has failed, thus mitigating erroneous switchover operations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If computing nodes share operational state information through internode communication link, then failover capability is improved, but false failure detection occurs due to communication link failures
Solution Approach 1:
The patent introduces a cloud environment node state provider as an intermediary component that mediates between computing nodes and the cloud infrastructure. This provider maintains authoritative state information about node operational status, allowing nodes to verify partner node status independently of direct communication link status, thus resolving the false failure detection problem while maintaining failover capability
Solution Approach 2:
The system implements feedback mechanisms where computing nodes continuously report their operational state to the cloud environment node state provider, which then provides this state information back to partner nodes. This feedback loop enables accurate failure detection by comparing expected state with actual reported state, distinguishing true failures from communication link issues
2Device complexity
If computing nodes rely on direct communication for state information, then system complexity is reduced, but service availability deteriorates during infrastructure failures
Solution Approach 1:
The cloud environment node state provider serves as an intermediary layer between computing nodes and cloud infrastructure, abstracting the complexity of infrastructure failure handling. Nodes query this provider for partner node status rather than directly monitoring infrastructure, maintaining simple node-to-node interactions while achieving high service availability through the mediator's reliable state information
3Speed
If computing nodes perform switchover operations on communication loss, then failover speed is improved, but erroneous switchover operations increase
Solution Approach 1:
The system performs preliminary verification of partner node failure status by querying the cloud environment node state provider before initiating switchover operations. This preliminary check confirms whether the partner node is truly failed or merely experiencing communication issues, preventing erroneous switchover while maintaining rapid failover capability when actual failures occur
Data Source
AI summary
One or more techniques and/or computing devices are provided for determining whether to perform a switchover operation between computing nodes. A first computing node and a second computing node, configured as disaster recovery partners, may be deployed within a computing environment. The first computing node and the second computing node may be configured to provide operational state information (e.g., normal operation, a failure, etc.) to a cloud environment node state provider and/or cloud persistent storage accessible through a cloud storage service. Accordingly, a computing node may obtain operational state information of a partner node from the cloud environment node state provider and/or the cloud storage service notwithstanding a loss of internode communication and/or an infrastructure failure that may otherwise appear as a failure of the partner node. In this way, the computing node may accurately determine whether the partner node has failed.


