IT Fault Root Cause Analysis via CMDB Event Correlation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current IT system fault analysis methods, such as those relying on user-defined causality settings, are time-consuming, dependent on technical knowledge, and inflexible in adapting to changes in system configuration, leading to instability and variability in fault root cause determination.
Innovation Solution
A computer-readable medium and method that analyzes root causes in IT systems by detecting incident configuration item events, identifying related events, generating event correlation information, and determining causality using a configuration management database, reducing user effort and enhancing analysis efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If user-defined causality setting method is used, then causality information can be obtained for fault analysis, but user effort and time consumption increase significantly
Solution Approach 1:
The system automatically discovers and defines causality relationships between configuration items by monitoring their operational states and interactions, eliminating the need for manual user definition. The fault analysis system self-configures the causality model based on observed system behavior and dependencies.
Solution Approach 2:
The manual mechanical process of user-defined causality setting is replaced with an automated information processing system that uses event correlation analysis and configuration data to dynamically determine causality relationships between system components.
2Reliability
If user-defined causality setting method is used, then causality can be established, but system stability decreases due to dependency on user knowledge
Solution Approach 1:
The system autonomously maintains causality relationships based on actual system configuration and operational data, ensuring consistency and stability. The causality model is continuously updated by the system itself rather than relying on external user knowledge that may vary in quality.
Solution Approach 2:
The system continuously monitors system operations and uses feedback from actual system behavior to validate and update causality relationships, ensuring they reflect the true system state and maintaining stability across different operational conditions.
3Ease of manufacture
If manual causality definition is required, then initial system design can be completed, but ease of operation deteriorates due to increased complexity
Solution Approach 1:
The system automatically generates the causality model from configuration item relationships and operational data, eliminating the complex manual setup process. Users simply need to deploy the system, which then self-configures the causality analysis framework.
Solution Approach 2:
The system performs preliminary automatic discovery and definition of causality relationships during initial deployment, so that when users need fault analysis capability, the causality model is already established and ready for use without requiring user intervention.
4Reliability
If user-defined causality is used, then fault analysis can be performed, but adaptability to system changes decreases
Solution Approach 1:
The causality model is dynamically updated based on real-time system configuration changes and operational data. When system configuration changes occur, the system automatically detects and updates the causality relationships to reflect the new system state, maintaining adaptability without requiring manual redefinition.
Solution Approach 2:
The system continuously monitors system configuration and operational state changes, using feedback to automatically update causality relationships. This ensures the fault analysis capability remains adapted to current system conditions while maintaining accurate root cause determination.
Data Source
AI summary
Disclosed is a management server for analyzing a root cause related with an abnormal situation in an IT system. This server has been made in an effort to analyze and present a root cause for a fault phenomenon which occurs in an IT system in order to satisfy a demand in the art. And the server has also been made in an effort to efficiently determine a potential fault related event in the IT system.


