Outage Prediction via Episode Segmentation in Enterprise Networks
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Enterprise network outages pose significant risks to revenue and reputation, and existing systems lack effective predictive capabilities to prevent or mitigate these outages, especially as complexity grows with customization and scale.
Innovation Solution
A system comprising a computer-readable medium, data collection engine, system behavior detection engine, outage prediction engine, and preventive operations engine, which uses univariate and multivariate anomaly detection, anomaly severity scaling, and episode tree data structures to predict potential outages and recommend proactive measures.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If enterprise network systems are customized and scaled to meet growing business needs, then system functionality and capacity are improved, but complexity increases resulting in massive amounts of data and a wide spectrum of behaviors that make outage prediction difficult
Solution Approach 1:
The patent segments the complex network monitoring data into distinct episodes representing different system states. Each episode is characterized by specific attributes (metrics, events, anomalies) that can be independently analyzed. This segmentation transforms the overwhelming mass of raw data into manageable, meaningful units that can be processed by machine learning models to predict outages without being overwhelmed by system complexity.
Solution Approach 2:
The patent introduces an intermediary layer of episode-based abstraction between the raw network data and the prediction system. Episodes serve as mediators that capture and structure the complex behaviors into standardized formats with defined attributes, enabling the machine learning models to process network data effectively without directly handling the full complexity of scaled enterprise systems.
2Reliability
If traditional monitoring systems are used to track network performance, then basic system status is maintained, but they lack predictive capabilities to resolve issues before outages occur
Solution Approach 1:
The patent implements preliminary action by training machine learning models on historical episode data to learn patterns that precede outages. The system performs preliminary analysis of current episodes against learned patterns to predict potential outages before they occur. This enables the system to take preventive actions—such as alerting operators or automatically triggering remediation workflows—before the actual outage impacts system availability.
Solution Approach 2:
The patent establishes a feedback loop where predicted outages and actual outage outcomes are continuously fed back into the machine learning models. This feedback mechanism allows the system to learn from past predictions and actual events, continuously improving its predictive accuracy. The feedback ensures the system maintains reliability by becoming progressively better at identifying precursors to outages, converting previously lost predictive information into actionable insights.
3Productivity
If reactive outage response is used, then immediate response to failures is provided, but downtime and revenue loss are maximized
Solution Approach 1:
The patent enables preliminary action by predicting outages before they occur, allowing operators to prepare remediation strategies and resources in advance. When an outage is predicted, the system can pre-stage recovery procedures, allocate necessary resources, and alert relevant personnel before the actual failure happens. This transforms the response model from reactive to proactive, significantly reducing actual downtime by eliminating the initial detection and preparation time that characterizes reactive responses.
Solution Approach 2:
The patent allows the system to skip the traditional outage detection phase by predicting failures before they manifest as actual outages. By identifying precursor patterns in episode data, the system rushes through the early stages of failure development that would otherwise lead to downtime. This skipping of the detection-and-response lag time directly reduces total downtime and associated revenue loss.
Data Source
AI summary
An effective strategy provides an intuitive starting point for an enterprise network agent to resolve issues before the issues increase the probability of an outage. Being able to predict whether and when a current anomalous state will transform into an outage is valuable to an enterprise network agent tasked with network administration, including monitoring the network; configuring the network; recommending software or hardware licenses, updates, or additions; obtaining software or hardware licenses or devices; generating reports and alerts; and launching countermeasures in association with the enterprise network.


