Managed Failover Service for High Availability
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing mechanisms for network-based failover services are overly complex, increase design work for customers, and lack features for customer visibility and control, leading to inadequate management of data integrity during failovers.
Innovation Solution
The implementation of a highly available failover service that coordinates failover workflows across multiple availability zones, allowing customers to manually or automatically trigger failovers, provides a visual editor for dependency tree creation, and ensures data integrity through event history logs and authoritative state information management.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If existing failover mechanisms are implemented, then failover capability is provided, but system complexity increases and customer control is reduced
Solution Approach 1:
The patent introduces a failover service as an intermediary layer between customer applications and the underlying infrastructure. This service abstracts the complex failover logic into managed workflows that customers can trigger with simple commands, reducing system complexity while maintaining reliability. The failover service coordinates state changes, manages dependencies, and executes reconciliation steps without requiring customers to implement complex failover mechanisms themselves.
Solution Approach 2:
The failover service enables self-service by allowing customers to manually trigger failovers through simple API calls or automated workflows without needing to understand or configure the underlying complex failover mechanisms. The service automatically manages the entire failover process including state transitions, dependency coordination, and data integrity reconciliation, making the system easy to use while maintaining high reliability.
2Ease of operation
If manual failover management is required, then customer control is increased, but operational overhead increases
Solution Approach 1:
The patent implements preliminary action by pre-defining failover workflows, dependency trees, and reconciliation steps before failures occur. Customers can configure their application dependencies and failover preferences in advance, so that when a failure occurs, the system can execute pre-planned failover sequences immediately. This reduces operational overhead while maintaining customer control over the failover behavior through advance configuration.
Solution Approach 2:
The failover service incorporates feedback mechanisms that provide customers with visibility into the failover process through event history logs, state information, and workflow progression tracking. Customers can monitor failover execution, view reconciliation status, and receive notifications about failover completion or issues. This feedback loop enables informed customer control without requiring continuous manual intervention, reducing operational overhead.
3Productivity
If failover workflows are automated, then operational efficiency is improved, but data integrity risks increase
Solution Approach 1:
The system performs preliminary reconciliation actions before completing failover transitions. Dependency trees are evaluated in advance to identify potential data integrity issues, and reconciliation workflows are prepared to address these issues before the failover is finalized. This preliminary validation ensures data integrity while maintaining automated operational efficiency.
Solution Approach 2:
The failover service implements feedback loops that continuously monitor data integrity during automated failover execution. Event history logs track state changes, and reconciliation workflows verify data consistency at critical transition points. If integrity issues are detected during automated failover, the system can pause or rollback the process, ensuring data integrity is maintained while preserving the benefits of automation.
4Loss of information
If comprehensive failover logging is implemented, then visibility is improved, but system overhead increases
Solution Approach 1:
The patent extracts logging and monitoring functions into a dedicated event history logging subsystem that is separate from the core failover execution logic. This extraction allows comprehensive logging of failover events, state changes, and workflow progression without burdening the main failover system. The logging subsystem captures necessary information for visibility while keeping the core failover mechanism clean and efficient, reducing overall system overhead.
Data Source
AI summary
The first computing system may interface with an operator of the application and a plurality of hosts of the application distributed between different partitions. The second and third computing systems may host first and second portion of the application in first and second partitions, respectively. The second and third computing systems may poll the first computing system to identify first and second value, respectively, representing state conditions of the first and second partitions, respectively, wherein the first and second partition state conditions are the active state, the passive state, and the fenced state. The second and third computing systems may receive responses from the first computing system comprising the first and second values, respectively, and based on the respective values, initiate a transition to the corresponding partition state condition. The first computing system may assign one of the first and second values to indicate which is the active state.


