Scalable Codebook Correlation for Cloud Topology Fault Identification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing systems for fault identification and alert management in complex distributed systems, such as cloud computing, face performance bottlenecks and increased complexity as network topology grows, leading to inefficient root cause analysis and alert correlation.
Innovation Solution
An adaptive fault identification system that retrieves alerts, correlates related symptoms using a codebook, and determines problem domains to identify root causes, eliminating duplicate symptoms and maintaining a scalable approach by dynamically computing causality maps and adapting to system changes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If a large hierarchical relationship structure of faults and alerts is maintained to provide comprehensive fault identification, then fault identification completeness is improved, but system complexity and computational requirements increase
Solution Approach 1:
The patent segments the large hierarchical relationship structure into multiple smaller codebooks organized in a tree structure. Each codebook represents a specific problem domain or subsystem, containing only the alerts and faults relevant to that domain. This segmentation reduces the complexity of individual codebooks while maintaining comprehensive fault identification coverage across the entire system through the hierarchical organization of multiple codebooks.
2Reliability
If the hierarchical relationship structure is expanded to cover cloud-scale topology, then fault identification coverage is improved, but performance bottlenecks and computational complexity worsen
Solution Approach 1:
The system divides the cloud-scale topology into multiple problem domains, each with its own codebook. This segmentation allows the system to process and correlate alerts within smaller, manageable codebooks rather than a single large hierarchical structure, significantly reducing computational complexity and improving performance while maintaining comprehensive coverage across the entire cloud-scale system.
Solution Approach 2:
The patent implements dynamic codebook generation and loading, where codebooks are created or loaded on-demand based on the specific alerts being processed. This dynamic approach allows the system to adapt to varying system states and topologies, loading only the necessary codebook segments when needed, thereby improving performance by avoiding the overhead of maintaining and processing the entire hierarchical structure continuously.
3Measurement precision
If comprehensive alert correlation is performed across the entire system repository, then root cause identification accuracy is improved, but computational complexity increases
Solution Approach 1:
The patent segments the alert correlation process into multiple stages, first filtering alerts to identify those belonging to the same problem domain using domain-specific codebooks, then performing detailed correlation only within those segmented groups. This segmented approach maintains high root cause identification accuracy by ensuring relevant alerts are correlated while significantly reducing computational complexity by avoiding unnecessary correlations across the entire system repository.
Solution Approach 2:
The patent introduces problem domains as intermediary layers between individual alerts and the root cause analysis process. These problem domains act as mediators that group related alerts together, enabling the correlation process to work on smaller, more manageable sets of alerts while still achieving comprehensive root cause identification. The intermediary problem domains reduce computational complexity by filtering and organizing alerts before detailed correlation analysis.
Data Source
AI summary
A system and algorithm to map alerts to a problem domain is provided so that the size of a codebook for the problem domain may be reduced and correlated independently of the general system and/or other problem domains in the system topology. When one or more symptoms of a fault appear in the system topology, the problem domain is discovered dynamically and the codebook for the problem domain generated dynamically. The system described herein provides for computation of a problem domain that has a reduced object repository for a set of objects which are directly or indirectly impacted by monitored symptoms. Multiple problem domains may be independently computed in order to build one or more codebooks. Each problem domain may be smaller compared to a system topology resulting in scale and performance improvements for codebook computation and correlation.


