Composite KPIs for Network Health Monitoring
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Manual monitoring of large volumes of telemetry data from network devices is time-consuming and cost-ineffective, especially in larger networks, making it challenging to detect anomalies and troubleshoot production networks efficiently.
Innovation Solution
A computer-implemented method using composite key performance indicators (KPIs) that combines multiple performance attributes of network devices to generate alerts, allowing for programmable alert logic and auto-remediation, and enabling operators to define composite KPIs and network health correlation logic via API calls, using flexible Boolean logic and YANG models for efficient health monitoring.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If manual monitoring of telemetry data is performed, then operators can detect anomalies and troubleshoot network issues, but the process becomes time-consuming and cost-ineffective as network size increases
Solution Approach 1:
The system enables self-service monitoring by automatically collecting telemetry data from network devices, processing it through analysis engines, and generating alerts without requiring manual operator intervention for each monitoring task. The automated alert generation and correlation systems perform monitoring functions independently, freeing operators from time-consuming manual monitoring while maintaining reliable anomaly detection.
Solution Approach 2:
The patent introduces intermediary components including telemetry analysis engines, alert correlation systems, and composite KPI generators that act as mediators between raw telemetry data and operator decision-making. These intermediaries automatically process and correlate data from multiple sources, reducing the time operators would otherwise spend manually analyzing telemetry data while improving detection reliability through systematic analysis.
2Reliability
If operators manually monitor multiple operational states simultaneously, then comprehensive network health monitoring is achieved, but the complexity and resource requirements increase significantly
Solution Approach 1:
The patent merges multiple monitoring functions into unified composite Key Performance Indicators (KPIs) that combine several operational states into single correlated alerts. The alert correlation system integrates data from multiple telemetry sources and operational states, presenting comprehensive monitoring coverage through consolidated views rather than requiring operators to monitor each state separately, thereby reducing system complexity while maintaining comprehensive coverage.
Solution Approach 2:
The monitoring system implements multi-functionality through universal alert correlation mechanisms that can handle multiple operational states, device types, and anomaly patterns through a single integrated framework. The composite KPI system provides universal monitoring capabilities across diverse network devices and states without requiring separate monitoring systems for each function, reducing overall system complexity while achieving comprehensive coverage.
3Measurement precision
If troubleshooting involves manual operations on many individual devices, then root cause analysis can be performed, but the process becomes time-consuming and cost-ineffective for larger networks
Solution Approach 1:
The patent introduces intermediary analysis systems including alert correlation engines and composite KPI generators that automatically perform root cause analysis by processing and correlating data from multiple devices. These intermediaries replace manual troubleshooting operations with automated analysis that maintains high measurement precision in identifying root causes while dramatically improving productivity by eliminating the need for operators to manually examine each individual device.
Solution Approach 2:
The system creates virtual copies of troubleshooting capabilities through automated analysis engines that replicate and execute diagnostic procedures across multiple devices simultaneously. Instead of requiring operators to physically or manually examine each device, the system uses software-based copies of diagnostic logic to perform parallel analysis, maintaining accurate root cause identification while significantly increasing troubleshooting efficiency through automation.
Data Source
AI summary
A remote server monitors the health of a network of computing devices through hierarchical composite indicators by obtaining performance attributes from computing devices in the network. The server generates a composite indicator associated with one or more of the computing device based on a combination of at least two performance attributes of the computing device(s). The server monitors the composite indicator and, responsive to a determination that the composite indicator indicates an alert condition, generates an alert associated with the computing device(s). Additionally, if the alert condition is subject to remediation, the server causes at least one of the computing devices to execute a command to provide remediation of the alert condition.


