Self-Healing Switch Network Event Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In complex communication networks, detecting and resolving anomalies in large-scale switch systems is inefficient due to the sheer size of log files and diversity of events, requiring manual inspection and error-prone administrative intervention, which hampers proactive issue resolution and self-healing capabilities.
Innovation Solution
An event analysis system within each switch uses pattern recognition and machine learning to identify events from log files, determine recovery actions, and execute self-healing processes, synchronizing across network instances to streamline anomaly detection and mitigation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If manual inspection of log files is used to detect network events, then administrators can identify issues, but the process becomes inefficient and time-consuming due to the large size of log files and diversity of events
Solution Approach 1:
The system enables self-service by automatically detecting, analyzing, and resolving network events without requiring manual administrator intervention. The event analysis system continuously monitors log files, identifies events using pattern recognition and machine learning, determines appropriate recovery actions, and executes them automatically, allowing the network to heal itself
Solution Approach 2:
The patent replaces the mechanical manual inspection process with an automated electronic system. Instead of administrators manually reviewing log files, the system uses machine learning models and pattern recognition algorithms to automatically analyze events, substitute human cognitive processes with computational algorithms that can process large volumes of data rapidly
2Reliability
If administrators intervene manually to resolve network events, then issues can be addressed, but the process becomes error-prone and hampers proactive issue resolution
Solution Approach 1:
The system implements feedback by continuously monitoring network events, analyzing their impact, evaluating recovery actions, and adjusting its behavior based on the outcomes. The event analysis system learns from past events and recovery outcomes, improving its ability to accurately identify issues and select appropriate remediation actions over time
Solution Approach 2:
The system enables self-service by automatically detecting, analyzing, and resolving network events without requiring manual administrator intervention. The event analysis system continuously monitors log files, identifies events using pattern recognition and machine learning, determines appropriate recovery actions, and executes them automatically, allowing the network to heal itself
3Adaptability or versatility
If the system processes all events in large-scale distributed networks, then comprehensive monitoring is achieved, but the complexity of analyzing diverse events across multiple switches increases
Solution Approach 1:
The system applies segmentation by dividing the network into individual switch instances, each with its own event analysis system that processes events locally. This distributed architecture allows each segment to independently analyze its own log files and events, reducing the complexity burden on any single system while maintaining comprehensive network-wide monitoring capability
Solution Approach 2:
The patent implements universality by creating a standardized event analysis system that can handle diverse event types across different switch models and network configurations. The machine learning models are trained on multiple event patterns and can generalize to handle various network events, making the system adaptable to different scenarios without requiring custom solutions for each event type
Data Source
AI summary
An event analysis system is provided. During operation, the system can determine an event description associated with the switch from an event log of the switch. The event description can correspond to an entry in a table in a switch configuration database of the switch. A respective database in the switch can be a relational database. The system can then obtain an event log segment, which is a portion of the event log, comprising the event description based on a range of entries. Subsequently, the system can apply a pattern recognition technique on the event log segment based on the entry in the switch configuration database to determine one or more patterns corresponding to an event associated with the event description. The switch can then apply a machine learning technique using the one or more patterns to determine a recovery action for mitigating the event.


