Master Softswitch Disaster Recovery Automation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current disaster recovery processes for wireless networks are inadequate as they do not efficiently identify the most synchronized replicate server to take over as the new master policy server, leading to manual and time-intensive emergency cutovers, which disrupt customer provisioning and network routing when the master policy server becomes unavailable.
Innovation Solution
A method to evaluate database synchronization between the master and replicate servers, identifying the best candidate to promote as the new master server, utilizing software scripting and automation to facilitate a seamless transition, including the use of AI components for classification and data transfer through APIs, GUIs, and CLIs.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If manual disaster recovery processes are used to identify and promote a replicate server as the new master policy server, then the process can be completed with existing systems, but the cutover becomes highly manual and time-intensive, disrupting customer provisioning and network routing
Solution Approach 1:
The system performs self-evaluation of replicate server synchronization status and automatically selects the best candidate to become the new master policy server without requiring manual intervention. The replicate servers autonomously report their database status, and the system automatically promotes the most synchronized replicate, eliminating manual assessment and reducing cutover time.
Solution Approach 2:
The system continuously monitors and evaluates the synchronization status of replicate servers before a disaster occurs. By maintaining up-to-date information about which replicates are most synchronized with the master database, the system is prepared to immediately identify the best candidate when a failure occurs, eliminating the need for time-consuming assessments during the actual cutover.
2Loss of time
If automated evaluation of replicate servers is implemented to identify the best candidate for promotion, then cutover time is reduced, but the system complexity increases due to additional evaluation mechanisms and AI components
Solution Approach 1:
The existing replicate server infrastructure is enhanced to perform multiple functions: maintaining database copies, reporting synchronization status, and participating in the selection process. The same communication channels and data structures used for normal database replication are leveraged for the evaluation process, avoiding the need for separate dedicated evaluation infrastructure.
Solution Approach 2:
Replicate servers continuously provide feedback about their synchronization status to the master policy server or a designated coordinator. This feedback mechanism uses existing database replication protocols and status reporting capabilities, allowing the system to gather necessary information through extended versions of current operations rather than requiring entirely new monitoring infrastructure.
3Reliability
If the most synchronized replicate server is automatically promoted to master role, then business continuity is ensured with minimal downtime, but manual control and oversight are reduced in the disaster recovery process
Solution Approach 1:
The system pre-evaluates and ranks replicate servers based on their synchronization status before a disaster occurs. This preliminary assessment creates a ready-made selection hierarchy that can be immediately acted upon when a failure occurs, ensuring the most suitable candidate is chosen without requiring complex real-time decision-making during the crisis.
Solution Approach 2:
The replicate servers autonomously determine their own synchronization status and communicate it to the system. The automated promotion process operates independently once triggered by a failure detection, reducing the need for human intervention while maintaining system reliability through algorithmic selection based on objective synchronization metrics.
Data Source
AI summary
A disaster recovery process and solution for a master routing server can utilize replicate servers to generate efficiencies when a failure occurs. A voice over Internet protocol (VoIP) platform can utilize the master routing server for customer call flow and feature provisioning. Back-up data can be stored at a cloud-based data store and sent to a replicate server that has been promoted in response to a failure being determined to have occurred. Additionally, the master routing server can receive and maintain a master provisioning database and send that database to the replicate routing servers for the purpose of properly routing call flows.


