Managed Failover Service for Application Availability
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing mechanisms for network-based failover services are overly complex, increase design work for customers, and lack features for customer visibility and control, leading to inadequate management of data integrity during failover events.
Innovation Solution
The implementation of a highly available managed failover service that coordinates failover workflows across multiple availability zones, allowing customers to manually or automatically trigger failovers, provides a visual editor for dependency tree creation, and ensures data integrity through authoritative state information management.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If existing mechanisms for network-based failover services are used, then failover functionality is provided, but the system complexity increases and customer design work increases
Solution Approach 1:
The patent introduces a failover service as an intermediary component that mediates between application partitions and availability zones. This service abstracts the complex failover logic into a manageable interface, reducing system complexity while maintaining reliability. The failover service coordinates state information and manages failover workflows without requiring customers to design complex failover mechanisms themselves.
Solution Approach 2:
The patent extracts failover functionality into a separate, dedicated failover service component. By taking out the complex failover logic from the customer's application design and placing it in a specialized service, the system reduces overall design complexity while preserving failover reliability. Customers only need to interact with simplified APIs rather than designing entire failover systems.
2Reliability
If existing failover mechanisms are used, then basic failover is achieved, but customer visibility and control are insufficient
Solution Approach 1:
The patent implements feedback mechanisms where the failover service provides state information about application partitions to customers. This visibility allows customers to monitor failover status and make informed decisions. The service feeds back operational states, health information, and failover progress, enabling customers to maintain control and visibility over their distributed applications during failover events.
Solution Approach 2:
The failover service enables customers to manually trigger failovers or configure automatic failover policies through simplified interfaces. Customers can self-manage failover operations without requiring deep technical knowledge of the underlying complexity. The service handles the complex coordination automatically while customers retain operational control through easy-to-use APIs and configurations.
3Manufacturing precision
If manual failover management is implemented, then data integrity can be maintained, but operational complexity increases
Solution Approach 1:
The failover service acts as an intermediary that manages data integrity during failover operations. It coordinates state information between partitions and ensures consistent data transitions without requiring customers to implement complex manual integrity checks. The service mediates read-write state transitions and fenced state management, preserving data integrity while reducing management complexity.
Solution Approach 2:
The patent implements preliminary actions by establishing fenced states and pre-configuring failover workflows before actual failover events occur. Application partitions are prepared in advance with defined state transitions and dependency relationships. This preliminary configuration ensures data integrity during failover without requiring complex real-time management decisions, as the failover service follows pre-established rules and state machines.
Data Source
AI summary
A computing system that receives and stores configuration information for the application in a data store. The configuration information comprises (1) identifiers for a plurality of cells of the application that include at least a primary cell and a secondary cell, (2) a defined state for each of the plurality of cells, (3) one or more dependencies for the application, and (4) a failover workflow defining actions to take in a failover event. The computing system receives an indication, from a customer, of a change in state of the primary cell or a request to initiate the failover event. The computing system updates, in the data store, the states for corresponding cells of the plurality of cells based on the failover workflow and updates, in the data store, the one or more dependencies for the application based on the failover workflow.


