Micro-Level Node Failover via Service KPI Monitoring
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Core networks typically monitor macro-level events and lack mechanisms to identify and address micro-level issues, leading to unnecessary service disruptions during node upgrades, as they fail to differentiate between operational and faulty services, resulting in premature node removal and prolonged service outages.
Innovation Solution
Implementing a failover and isolation server (FIS) system that monitors service-specific Key Performance Indicators (KPIs) to identify specific services causing outages and reroute affected requests to redundant nodes, minimizing service disruptions by maintaining node operationality and allowing selective service rerouting.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If core networks monitor only macro-level events, then the monitoring system remains simple, but micro-level service outages cannot be identified and addressed
Solution Approach 1:
The patent segments the monitoring system into macro-level event monitoring (existing simple infrastructure) and micro-level service KPI monitoring (new detailed tracking). By dividing monitoring into these separate layers, the system achieves precise micro-level detection without completely redesigning the entire monitoring infrastructure, thus resolving the contradiction between measurement precision and device complexity.
Solution Approach 2:
The patent introduces service KPIs as intermediary metrics that bridge the gap between macro-level node monitoring and micro-level service monitoring. These KPIs (such as call completion rates, data session success rates) act as mediators that translate service-level issues into measurable parameters, enabling precise detection without requiring direct complex monitoring of every service interaction.
2Reliability
If the entire node is removed from service during upgrades, then service outages are avoided, but network productivity decreases due to unnecessary service disruptions
Solution Approach 1:
The patent segments node services into individual service components that can be monitored and isolated independently. Instead of removing the entire node from service during upgrades, the system identifies and isolates only the specific faulty services, allowing other services to continue operating. This segmentation resolves the contradiction by maintaining service continuity (reliability) while preserving network throughput (productivity).
Solution Approach 2:
The patent applies local quality by treating different services on a node differently based on their health status. Healthy services continue to operate with normal quality, while faulty services are isolated with reduced or zero quality. This localized approach ensures that service disruptions are minimized to only the affected areas, maintaining overall network productivity while ensuring reliability for functional services.
3Ease of repair
If service-specific monitoring is implemented, then faulty services can be identified and isolated, but device complexity increases
Solution Approach 1:
The patent implements feedback mechanisms where service KPIs continuously monitor service health and automatically trigger isolation actions when thresholds are breached. This automated feedback loop simplifies the complexity by removing the need for manual analysis and decision-making, allowing the system to self-diagnose and self-isolate faulty services with minimal human intervention, thus improving ease of repair while managing complexity.
Solution Approach 2:
The patent enables the monitoring system to perform self-service by automatically identifying and isolating faulty services based on KPI thresholds without requiring manual operator intervention. The system self-diagnoses service health, determines isolation necessity, and executes isolation actions autonomously. This self-service capability improves ease of repair while containing complexity within automated processes rather than requiring complex manual procedures.
4Reliability
If nodes are removed from service during upgrades, then service outages are prevented, but service outages are prolonged due to rollbacks
Solution Approach 1:
The patent applies preliminary action by monitoring service KPIs before complete service failure occurs and isolating services proactively. Instead of waiting for macro-level failures to trigger node removal, the system detects deteriorating service conditions through KPI trends and isolates affected services early in the degradation process. This preliminary action prevents complete service outages and eliminates the need for time-consuming rollbacks, thus improving service availability while minimizing loss of time.
Data Source
AI summary
An improved core network that can monitor micro-level issues, identify specific services of specific nodes that may be causing an outage, and perform targeted node failovers in a manner that does not cause unnecessary disruptions in service is described herein. For example, the improved core network can include a failover and isolation server (FIS) system. The FIS system can obtain service-specific KPIs from the various nodes in the core network. The FIS can then compare the obtained KPI values of the respective service with corresponding threshold values. If any KPI value exceeds a corresponding threshold value, the FIS may preliminarily determine that the service of the node associated with the KPI value is responsible for a service outage. The FIS can initiate a failover operation, which causes the node to re-route any received requests corresponding to the service potentially responsible for the service outage to a redundant node.


