Automated Secondary Site Failover Module for Messaging Platforms
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional messaging platforms lack automated failover support, leading to customer impacts during system failures in primary or secondary data centers, as they require complex settings and additional protections to prevent circular messaging, making them ineffective in automatically switching replication flows during failures.
Innovation Solution
An automated secondary site failover module that establishes a communication link between data center sites, monitors database states, and automatically switches replication flows from a passive to an active state, ensuring seamless data replication and minimizing customer impact during failures.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If conventional messaging platforms are used for data center mirroring, then data replication can be established, but automated failover support is lacking due to complex settings and additional protections required to prevent circular messaging
Solution Approach 1:
The patent introduces a failover manager as an intermediary component that sits between the messaging platforms and coordinates failover operations. This manager handles the complex logic of detecting failures, determining failover eligibility, and executing the switch, thereby isolating the complexity from the core messaging platform and enabling automated failover without requiring changes to the messaging platform itself.
Solution Approach 2:
The system implements feedback mechanisms where the failover manager continuously monitors the health status of primary and secondary data centers, receives updates on replication status, and uses this information to automatically trigger failover when needed. The feedback loop includes health checks, replication status monitoring, and automatic decision-making based on predefined criteria, eliminating the need for manual intervention.
2Productivity
If manual failover management is implemented, then system control is maintained, but customer impact increases during system failures due to lack of automatic resolution
Solution Approach 1:
The system performs preliminary actions by pre-configuring failover policies, eligibility criteria, and recovery procedures before failures occur. The failover manager is pre-loaded with the necessary logic and parameters to automatically execute failover decisions, eliminating delays associated with manual assessment and configuration during actual failure events.
Solution Approach 2:
The failover system operates autonomously by self-monitoring data center health, self-deciding when failover is necessary based on predefined criteria, and self-executing the failover process. This self-service capability eliminates the need for human operators to intervene during failures, thereby reducing customer impact time and enabling immediate automated recovery.
3Reliability
If data replication is maintained between primary and secondary sites, then data availability is improved, but system complexity increases requiring state monitoring and switching mechanisms
Solution Approach 1:
The system segments the failover management functionality into distinct modular components: health monitoring modules, replication status tracking modules, decision-making logic, and execution modules. This segmentation allows each component to handle specific aspects of the complex failover process independently, making the overall system more manageable and easier to maintain while ensuring data availability.
Data Source
AI summary
Automated disaster recovery site failover of a messaging platform is disclosed. A processor establishes a communication link between a first data center site and a second data center site via a communication network. The first data center site includes a first database in an active state and the second data center site includes a second database in a passive state during which data replication flows from the first database to the second database. The processor monitors states of the first database and the second database; detects, in response to monitoring, that the first database has changed its state from the active state to the passive state and that the second database has changed its state from the passive state to the active state; and automatically switches, in response to detecting, the data replication flows during which the data replication flows from the second database to the first database.


