Centrality Algorithms for Root Cause Identification in Computing Systems
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for determining the root cause of data events in computing systems are inefficient and often resource-intensive, leading to wasted resources and prolonged downtime, as they fail to accurately identify the actual cause of the issue, often targeting affected services rather than the root cause.
Innovation Solution
The use of specially configured centrality algorithms to analyze a directed dependency graph of computing system services, identifying the root cause service and generating a prioritized list to focus maintenance efforts on the most likely cause, thereby reducing unnecessary resource allocation and improving efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If current methods are used to determine root cause of data events, then comprehensive analysis of affected services is performed, but resource efficiency deteriorates and identification accuracy worsens due to targeting affected services rather than actual root cause
Solution Approach 1:
The patent segments the computing system into a directed dependency graph where services are nodes and dependencies are edges. This segmentation allows the system to analyze only the relevant portion of the system affected by a data event, rather than performing comprehensive analysis of all services. The affected services subgraph is extracted and analyzed separately, improving both accuracy and resource efficiency.
Solution Approach 2:
The patent performs preliminary actions by pre-configuring centrality algorithms and pre-establishing the directed dependency graph structure before data events occur. When a data event is detected, the system can immediately apply the pre-configured algorithms to the pre-built graph structure, avoiding the need to perform comprehensive analysis from scratch and thereby improving resource efficiency.
2Measurement precision
If comprehensive analysis of all affected services is performed, then thorough investigation is achieved, but downtime increases due to prolonged analysis time
Solution Approach 1:
The patent extracts the affected services subgraph from the complete directed dependency graph, isolating only the relevant portion of the system that needs analysis. This extraction eliminates the need to analyze unrelated services, significantly reducing analysis time while maintaining identification accuracy by focusing computational resources on the minimal necessary subset of services.
Solution Approach 2:
The patent segments the analysis scope into an affected services subgraph, separating it from the rest of the computing system. This segmentation allows parallel processing and focused analysis on only the relevant services, reducing overall analysis time while maintaining thoroughness within the segmented scope.
3Reliability
If maintenance efforts are distributed across all affected services, then comprehensive coverage is achieved, but productivity decreases due to wasted resource allocation
Solution Approach 1:
The patent applies local quality by prioritizing maintenance efforts based on the likelihood of each service being the root cause. Instead of uniform maintenance distribution, the system assigns different priority levels to different services based on centrality algorithm results, directing resources to the most critical services first. This improves maintenance efficiency while ensuring system reliability through focused intervention.
Solution Approach 2:
The patent changes the parameter of maintenance priority from uniform to variable, based on the output of centrality algorithms. Services are assigned priority levels according to their calculated likelihood of being the root cause, transforming the maintenance strategy from broad coverage to targeted intervention, thereby improving productivity without compromising reliability.
Data Source
AI summary
Embodiments of the present disclosure provide improved identification and handling of root causes for data event(s). Some embodiments improve the accuracy of determinations of a root cause or likely order of root causes of a data event affecting any number of system(s), and cause transmission of data associated with such root cause(s) for use in triaging such data event(s) and/or facilitating efficient servicing to resolve the data event. Some embodiments utilize modified centrality algorithm(s) to efficiently and accurately identify a likely root cause of a data event in a computing environment. Some embodiments generate and/or output notifications that indicate the particular computing system(s) identified as a root cause of a data event, and/or the particular computing system(s) identified not as a root cause but affected by a data event of the root cause computing system.


