Automated Regional Failover Using Service Weight Modification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing manual and sequential failover processes for switching services from a primary region to a secondary region in case of a fault are inefficient and unreliable, leading to prolonged downtime and reduced availability of services.

Innovation Solution

An automated method that prepares an input file with a list of service names, modifies service weights, and introduces a sleep time to simultaneously fail over all services and their corresponding databases from the primary region to the secondary region upon detection of a fault.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If manual and sequential failover processes are used to switch services from primary region to secondary region, then operational control and safety are maintained, but failover time increases and service availability decreases

Engineering Contradiction:
Improveservice availabilityVSAvoidfailover time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system pre-preares an input file containing the list of all services operating on the primary region before a fault occurs. This preliminary preparation enables the failover process to immediately execute with `kubectl failover` without manual intervention, significantly reducing failover time while maintaining reliability through automated execution of pre-planned service migration.

Inventive Principle:
Principle #10Preliminary action

2Productivity

If sequential failover of each service is implemented, then control over each service migration is maintained, but overall failover duration increases

Engineering Contradiction:
Improvefailover speedVSAvoidtotal failover duration
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent merges the failover operations of multiple services into a single simultaneous execution using `kubectl failover` with the `--all` flag. This combines individual service failover commands into one unified operation that migrates all services from the primary region to the secondary region concurrently, dramatically increasing failover speed while the input file ensures all necessary service information is ready in advance.

Inventive Principle:
Principle #5Merging (Combining)

3Productivity

If automated failover without service list preparation is used, then failover speed increases, but system complexity and potential errors increase

Engineering Contradiction:
Improvefailover execution speedVSAvoidsystem configuration complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system generates and stores an input file containing the complete list of services operating on the primary region in advance. This preliminary action simplifies the automated failover process by providing a ready-to-use service list that `kubectl failover` can immediately process, reducing system complexity while maintaining high failover execution speed through automation.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12339753B2Automated regional failover
Publication Date: 2025.06.24 CAPITAL ONE SERVICES LLC
  • US12339753B2 patent drawing
  • US12339753B2 patent drawing
  • US12339753B2 patent drawing

AI summary

Disclosed herein are system, method, and computer program product embodiments for automatically failing over all services operating on a primary region to a secondary region upon detection or notification of a fault in the primary region. When a fault exists on the primary region, the method traverses each cluster containing services operating on the primary region and prepares an input file including a list of service names identifying each service operating on the primary region. Referencing the input file, the method fails over each service from the primary region to the secondary region by modifying a service weight corresponding to each service. This failover process of services may be done simultaneously with failing over any databases corresponding to the failed-over services from the primary region to the secondary region. The method may also introduce a sleep time after modifying each service weight to avoid any potential throttling issues.