Managed Failover Service for High Availability

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing mechanisms for network-based failover services are overly complex, increase design work for customers, and lack features for customer visibility and control, leading to inadequate management of data integrity during failovers.

Innovation Solution

The implementation of a highly available failover service that coordinates failover workflows across multiple availability zones, allowing customers to manually or automatically trigger failovers, provides a visual editor for dependency tree creation, and ensures data integrity through event history logs and authoritative state information management.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If existing failover mechanisms are implemented, then failover capability is provided, but system complexity increases and customer control is reduced

Engineering Contradiction:
Improvefailover capabilityVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent introduces a failover service as an intermediary layer between customer applications and the underlying infrastructure. This service abstracts the complex failover logic into managed workflows that customers can trigger with simple commands, reducing system complexity while maintaining reliability. The failover service coordinates state changes, manages dependencies, and executes reconciliation steps without requiring customers to implement complex failover mechanisms themselves.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The failover service enables self-service by allowing customers to manually trigger failovers through simple API calls or automated workflows without needing to understand or configure the underlying complex failover mechanisms. The service automatically manages the entire failover process including state transitions, dependency coordination, and data integrity reconciliation, making the system easy to use while maintaining high reliability.

Inventive Principle:
Principle #25Self-service

2Ease of operation

If manual failover management is required, then customer control is increased, but operational overhead increases

Engineering Contradiction:
Improvecustomer controlVSAvoidoperational overhead
Core Design Contradiction:
Ease of operationVSLoss of time

Solution Approach 1:

The patent implements preliminary action by pre-defining failover workflows, dependency trees, and reconciliation steps before failures occur. Customers can configure their application dependencies and failover preferences in advance, so that when a failure occurs, the system can execute pre-planned failover sequences immediately. This reduces operational overhead while maintaining customer control over the failover behavior through advance configuration.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The failover service incorporates feedback mechanisms that provide customers with visibility into the failover process through event history logs, state information, and workflow progression tracking. Customers can monitor failover execution, view reconciliation status, and receive notifications about failover completion or issues. This feedback loop enables informed customer control without requiring continuous manual intervention, reducing operational overhead.

Inventive Principle:
Principle #23Feedback

3Productivity

If failover workflows are automated, then operational efficiency is improved, but data integrity risks increase

Engineering Contradiction:
Improveoperational efficiencyVSAvoiddata integrity
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The system performs preliminary reconciliation actions before completing failover transitions. Dependency trees are evaluated in advance to identify potential data integrity issues, and reconciliation workflows are prepared to address these issues before the failover is finalized. This preliminary validation ensures data integrity while maintaining automated operational efficiency.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The failover service implements feedback loops that continuously monitor data integrity during automated failover execution. Event history logs track state changes, and reconciliation workflows verify data consistency at critical transition points. If integrity issues are detected during automated failover, the system can pause or rollback the process, ensuring data integrity is maintained while preserving the benefits of automation.

Inventive Principle:
Principle #23Feedback

4Loss of information

If comprehensive failover logging is implemented, then visibility is improved, but system overhead increases

Engineering Contradiction:
ImprovevisibilityVSAvoidsystem overhead
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent extracts logging and monitoring functions into a dedicated event history logging subsystem that is separate from the core failover execution logic. This extraction allows comprehensive logging of failover events, state changes, and workflow progression without burdening the main failover system. The logging subsystem captures necessary information for visibility while keeping the core failover mechanism clean and efficient, reducing overall system overhead.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS11366728B2Systems and methods for enabling a highly available managed failover service
Publication Date: 2022.06.21 AMAZON TECH INC
  • US11366728B2 patent drawing
  • US11366728B2 patent drawing
  • US11366728B2 patent drawing

AI summary

The first computing system may interface with an operator of the application and a plurality of hosts of the application distributed between different partitions. The second and third computing systems may host first and second portion of the application in first and second partitions, respectively. The second and third computing systems may poll the first computing system to identify first and second value, respectively, representing state conditions of the first and second partitions, respectively, wherein the first and second partition state conditions are the active state, the passive state, and the fenced state. The second and third computing systems may receive responses from the first computing system comprising the first and second values, respectively, and based on the respective values, initiate a transition to the corresponding partition state condition. The first computing system may assign one of the first and second values to indicate which is the active state.