Telecom Peer Device Health Monitoring for Controlled Failover
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing telecommunications networks face challenges in maintaining high availability with minimal impact during device switchovers and lack user control over failover decisions, as traditional redundancy mechanisms fail to account for deterministic behaviors and criticality of failures.
Innovation Solution
A high availability configuration with two peer device platforms, each equipped with a device health monitoring component and a rules engine, generates health counts based on detected failures and predetermined rules, allowing for informed failover decisions that consider criticality and redundancy, thereby reducing network impact and enhancing user control.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional redundancy mechanisms are used for failover, then high availability is provided, but user control over failover decisions is lacking and network impact is not minimized
Solution Approach 1:
The system implements feedback mechanisms where health monitoring components continuously assess device status and feed this information to rules engines. The rules engines process this feedback against predetermined rules and generate health counts, which are then used to make informed failover decisions. This closed-loop feedback system enables both automatic reliability maintenance and user-controlled failover based on assessed health conditions.
Solution Approach 2:
The system dynamically adjusts failover behavior by continuously monitoring device health and recalculating health counts based on current conditions. Rather than static redundancy configurations, the system adapts its failover decisions in real-time based on changing device states, allowing users to control failover based on current system conditions while maintaining high availability.
2Reliability
If traditional redundancy mechanisms perform switchovers, then failover is achieved, but minimal network impact is not ensured
Solution Approach 1:
The system performs preliminary health assessments and generates health counts before failover decisions are made. By pre-evaluating device health conditions and predicting potential network impact, the system can make informed decisions about whether to proceed with failover. This preliminary action prevents unnecessary switchovers that would cause network disruption while ensuring failover occurs when truly needed.
Solution Approach 2:
The system replaces traditional mechanical or direct failover mechanisms with a rules-based decision engine that processes health information and determines optimal failover timing. This substitution allows for intelligent evaluation of network impact before executing failover, minimizing disruption by only switching when necessary and using the least disruptive method available.
3Device complexity
If deterministic behaviors are not considered in failover decisions, then simple redundancy switching is achieved, but accurate failover decision making is compromised
Solution Approach 1:
The system changes the parameters used in failover decisions from simple binary failure states to multi-dimensional health counts that incorporate deterministic device behaviors. By tracking multiple health parameters and weighting them according to predetermined rules, the system achieves accurate failure assessment while maintaining manageable complexity through standardized parameter changes.
Solution Approach 2:
The rules engine serves multiple functions: it monitors device health, applies deterministic behavior rules, generates health counts, and makes failover decisions. This universal component handles diverse assessment requirements within a single framework, achieving accurate failure assessment without proportionally increasing system complexity by consolidating multiple functions into one multi-functional engine.
Data Source
AI summary
Systems and methods of providing high availability of telecommunications systems and devices in a telecommunications network. A telecommunications device is deployed in a high availability configuration that includes two or more peer device platforms, in which each peer device platform can operate in either an active mode or a standby mode. Each peer device platform includes a device health monitoring component and a rules engine. By detecting one or more failures and/or faults associated with the peer device platforms using the respective device health monitoring components, and generating, using the rules engine, a health count for each peer device platform based on the detected failures/faults and one or more predetermined rules, failover decisions can be made based on a comparison of the health counts for the respective peer device platforms, while reducing the impact on the telecommunications network and providing an increased level of user control over the failover decisions.


