Failover Service for Block Storage Volumes

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing mechanisms for managing failovers in network-based computing are overly complicated, inefficient, and lack customer visibility and control, leading to increased design work and potential data integrity issues during failures.

Innovation Solution

A system for managing network-based failover services that coordinates failover workflow design and execution, ensuring predictable and controlled failover by replicating routing criteria between network devices and enabling automatic routing of traffic to secondary regions in case of failure.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If existing failover mechanisms are used, then failover capability is provided, but the system becomes overly complicated and inefficient

Engineering Contradiction:
Improvefailover capabilityVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent extracts the failover management functionality into a dedicated failover service that operates independently from the primary storage system. This service handles failover coordination, routing criteria replication, and traffic redirection, separating these complex operations from the core storage operations to reduce overall system complexity while maintaining reliability.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The failover service acts as an intermediary between the primary storage system and the secondary storage system. It receives failure notifications, determines failover conditions, replicates routing criteria to network devices, and coordinates traffic redirection, thereby simplifying the interaction between multiple systems during failover events.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If existing failover mechanisms are used, then failover capability is provided, but customer visibility and control are lacking

Engineering Contradiction:
Improvefailover capabilityVSAvoidcustomer visibility and control
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The failover service implements feedback mechanisms by monitoring storage system health status, detecting failure conditions, and providing controlled notifications to customers. It maintains visibility into failover progress and allows customers to understand the failover process through structured communication while preserving system automation.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system provides dynamic control over failover processes, allowing customers to configure failover policies, select target storage systems, and control the failover timing. This dynamic configuration capability gives customers operational control while the system automatically executes the failover when conditions are met.

Inventive Principle:
Principle #15Dynamics

3Ease of operation

If manual failover processes are used, then customer control is maintained, but design work and operational burden increase

Engineering Contradiction:
Improvecustomer controlVSAvoiddesign work and operational burden
Core Design Contradiction:
Ease of operationVSLoss of time

Solution Approach 1:

The failover service performs preliminary actions by pre-configuring failover policies, pre-establishing routing criteria, and pre-validating secondary storage systems before failures occur. This preparation eliminates the need for manual design work during actual failover events, reducing operational burden while maintaining customer control through configurable policies.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The failover service enables self-service failover execution where the system automatically detects failures, determines appropriate target systems, replicates routing criteria, and redirects traffic without requiring manual intervention. This automation reduces operational burden while customers retain control through policy configuration and monitoring capabilities.

Inventive Principle:
Principle #25Self-service

4Extent of automation

If failover is not automated, then customer involvement is minimized, but data integrity issues may occur during failures

Engineering Contradiction:
Improvefailover automationVSAvoiddata integrity
Core Design Contradiction:
Extent of automationVSReliability

Solution Approach 1:

The failover service implements beforehand cushioning by continuously monitoring storage system health, maintaining backup routing criteria in secondary systems, and preparing failover sequences in advance. This preparation ensures that when failures occur, the system can immediately redirect traffic to healthy systems without data loss, cushioning against data integrity issues through proactive preparation.

Inventive Principle:
Principle #11Beforehand cushioning (Prior cushioning)

Solution Approach 2:

The system uses feedback mechanisms to continuously monitor data integrity during failover processes, verifying that routing criteria are correctly replicated and that traffic redirection maintains data consistency. This feedback loop ensures data integrity is maintained throughout the automated failover process by detecting and preventing integrity violations.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS12271276B1Systems and methods for enabling a failover service for block-storage volumes
Publication Date: 2025.04.08 AMAZON TECH INC
  • US12271276B1 patent drawing
  • US12271276B1 patent drawing
  • US12271276B1 patent drawing

AI summary

The present disclosure generally relates to a first network device in a primary region that can failover network traffic into a second network device in a failover region. The first network device can receive routing criteria identifying how traffic originating in the primary region should be routed. The first network device can transmit this routing criteria to the second network device in the failover region. Based on determining the occurrence of a failover event, the first network device may transmit network traffic originating in the primary region to the second network device in the failover region. The second network device can determine how to route the network traffic based on the routing criteria of the primary region. In some embodiments, the second network device can determine how to route the network traffic based on the routing criteria of the failover region.