Machine-Learned Service Maps for Faster Network Root Cause Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current root cause analysis techniques in networks are inaccurate and time-consuming, failing to identify the full extent of impacted systems, devices, applications, and services, leading to prolonged device and service unavailability.
Innovation Solution
Implementing machine-learning based service mapping to rapidly identify candidate service maps, determining relationships between configuration items using databases, log data, and network traffic, and using these maps to enhance event management for faster and more accurate root cause analysis.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional root cause analysis techniques are used, then the analysis process is simple, but the accuracy of identifying impacted systems and services is low and the analysis time is long
Solution Approach 1:
The system performs preliminary service mapping and relationship discovery before incidents occur, building a knowledge base of service dependencies and configurations. This pre-established context enables rapid root cause identification when incidents occur, eliminating the need for time-consuming analysis during actual problems
Solution Approach 2:
The system creates virtual representations (service maps) of the actual network infrastructure, capturing relationships between configuration items, services, and dependencies. These copied models can be queried and analyzed without affecting the real system, enabling fast root cause determination through map traversal rather than live system analysis
2Loss of information
If traditional event management is used, then the system complexity is low, but the ability to identify the full extent of impacted services is insufficient
Solution Approach 1:
The system segments the network infrastructure into discrete configuration items and services, establishing explicit relationships between them in a service map. This segmentation allows comprehensive tracking of which specific services and components are impacted by incidents, providing complete visibility into the extent of service degradation
Solution Approach 2:
The service map acts as an intermediary layer between raw event data and impact analysis. This intermediate representation captures the complex relationships between configuration items and services, enabling comprehensive impact assessment without requiring direct complex queries to the underlying infrastructure
Data Source
AI summary
An implementation may involve: obtaining a representation of a network event relating to a network, wherein the network enables operation of a plurality of services each involving one or more computing devices or software applications; obtaining information associated with the network event, wherein the information identifies one of the computing devices or the software applications; based on the information, identifying a subset of services of the plurality of services based on determining that each of the subset of services satisfies an impact criterion with respect to the network event, wherein the subset of services are associated with candidate service maps that were generated by a machine learning process; and providing an indication that the subset of services are related to the network event.


