Machine Learning Internet Outage Detection from End-User Telemetry
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for detecting internet outages are ineffective as they rely on self-reporting by users and fail to accurately estimate the scale and magnitude of outages, making it difficult to understand their impact on users and network infrastructure.
Innovation Solution
A user performance monitoring solution that collects telemetry data from end-user devices to create performance baselines and uses machine learning models to detect and quantify internet outages, identifying blackouts and brownouts by analyzing metrics such as Page Fetch Time, Time To First Byte, probe error rates, latency, and packet drops, and visualizing the outage magnitude.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If self-reporting by users through social media is used to detect outages, then outage detection is possible, but the scale and magnitude of outages cannot be accurately estimated
Solution Approach 1:
The patent introduces specialized monitoring sites and telemetry collection systems as intermediaries between users and the detection system. These intermediaries automatically collect performance metrics (Page Fetch Time, Time to First Byte, probe error rates, latency, packet drops) from multiple sources including social media mentions and dedicated monitoring infrastructure, transforming unstructured user reports into quantifiable data that enables accurate outage magnitude estimation without requiring direct user reporting
Solution Approach 2:
The patent replaces the manual self-reporting mechanism with automated machine learning-based detection systems. The ML model processes telemetry data from multiple sources to automatically identify and characterize outages, substituting the mechanical process of user reporting and manual analysis with computational automation that provides both detection and quantitative assessment simultaneously
2Measurement precision
If telemetry data is collected from millions of devices across all ISPs, then accurate outage detection is achieved, but data processing complexity increases
Solution Approach 1:
The patent segments the vast telemetry data from millions of devices by organizing it according to ISP tiers and geographic regions. The system divides the monitoring landscape into manageable segments (different ISP levels, regional networks) and applies specialized analysis to each segment, allowing accurate detection across the entire internet infrastructure while processing data in organized, tractable units rather than as an overwhelming monolith
Solution Approach 2:
The patent creates a universal ML-based detection framework that processes diverse telemetry data types (Page Fetch Time, TTFB, error rates, latency, packet drops) from multiple sources (user devices, specialized monitoring sites, social media) through a single unified system. This multi-functional approach handles varied data formats and sources consistently, reducing processing complexity by applying the same analytical framework across all data types rather than requiring separate systems for each
3Measurement precision
If machine learning models are trained on multiple performance metrics, then blackout and brownout prediction accuracy improves, but model training time and computational resources increase
Solution Approach 1:
The patent performs preliminary actions by pre-training the machine learning model on historical telemetry data encompassing multiple performance metrics (Page Fetch Time, Time to First Byte, probe error rates, latency, packet drops) and various outage scenarios. This pre-training establishes baseline patterns and relationships in the data before deployment, so that when the model is deployed, it can quickly assess current conditions against pre-learned patterns without requiring extensive real-time computation or retraining, thus achieving high prediction accuracy with minimal operational training time
Data Source
AI summary
The present systems and methods provide a user performance monitoring solution that enables the monitoring of application and device performance from the end user's point of view. The present systems and methods help Information Technology (IT) personnel to ensure the quality of digital experience across the enterprise. The present system is adapted to collect telemetry data from devices relative to the performance of all tiers of Internet Service Providers (ISPs), create a baseline of the performance of the ISPs based on a plurality of metrics and the collected telemetry data, train a Machine Learning (ML) model to assess blackout and brownout prediction accuracy at different performance values for the metrics, and identify a blackout or brownout, wherein a blackout or brownout is identified when real time performance is worse than the performance values identified by the model.


