Automated Regional Failover Using Service Weight Modification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing manual and sequential failover processes for switching services from a primary region to a secondary region in case of a fault are inefficient and unreliable, leading to prolonged downtime and reduced availability of services.
Innovation Solution
An automated method that prepares an input file with a list of service names, modifies service weights, and introduces a sleep time to simultaneously fail over all services and their corresponding databases from the primary region to the secondary region upon detection of a fault.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If manual and sequential failover processes are used to switch services from primary region to secondary region, then operational control and safety are maintained, but failover time increases and service availability decreases
Solution Approach 1:
The system pre-preares an input file containing the list of all services operating on the primary region before a fault occurs. This preliminary preparation enables the failover process to immediately execute with `kubectl failover` without manual intervention, significantly reducing failover time while maintaining reliability through automated execution of pre-planned service migration.
2Productivity
If sequential failover of each service is implemented, then control over each service migration is maintained, but overall failover duration increases
Solution Approach 1:
The patent merges the failover operations of multiple services into a single simultaneous execution using `kubectl failover` with the `--all` flag. This combines individual service failover commands into one unified operation that migrates all services from the primary region to the secondary region concurrently, dramatically increasing failover speed while the input file ensures all necessary service information is ready in advance.
3Productivity
If automated failover without service list preparation is used, then failover speed increases, but system complexity and potential errors increase
Solution Approach 1:
The system generates and stores an input file containing the complete list of services operating on the primary region in advance. This preliminary action simplifies the automated failover process by providing a ready-to-use service list that `kubectl failover` can immediately process, reducing system complexity while maintaining high failover execution speed through automation.
Data Source
AI summary
Disclosed herein are system, method, and computer program product embodiments for automatically failing over all services operating on a primary region to a secondary region upon detection or notification of a fault in the primary region. When a fault exists on the primary region, the method traverses each cluster containing services operating on the primary region and prepares an input file including a list of service names identifying each service operating on the primary region. Referencing the input file, the method fails over each service from the primary region to the secondary region by modifying a service weight corresponding to each service. This failover process of services may be done simultaneously with failing over any databases corresponding to the failed-over services from the primary region to the secondary region. The method may also introduce a sleep time after modifying each service weight to avoid any potential throttling issues.


