IT System Fault Event Detection via Statistical Window Slicing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current techniques for root cause analysis in IT systems are time-consuming, require significant user effort, and are dependent on user expertise, making them inefficient and unstable, especially when system configurations change.
Innovation Solution
A computer program that collects performance data, generates statistical representative values using window slicing and machine learning algorithms, and determines potential events in IT systems, reducing the need for user-defined causality settings and improving system stability by automating the detection of abnormal situations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If user-defined causality setting method is used for root cause analysis, then causality information can be obtained, but user effort and time consumption increase significantly
Solution Approach 1:
The system automatically collects performance data from multiple entities, generates performance information data, and determines causality relationships without requiring user definition. The causality determination module autonomously analyzes the performance information data to establish causality models, enabling the system to serve itself rather than relying on user input for causality settings.
Solution Approach 2:
The system transforms raw performance data into performance information data through statistical processing, changing the parameter representation from raw metrics to derived information that reveals causality relationships. This parameter transformation enables automatic causality determination without user intervention.
2Reliability
If user-defined causality setting method is used, then causality can be established, but dependency on user expertise increases
Solution Approach 1:
The causality determination module automatically analyzes performance information data to establish causality relationships between entities without requiring user expertise. The system independently performs statistical analysis and determines causality models, eliminating dependency on user knowledge and experience.
Solution Approach 2:
The system replaces the manual mechanical process of user-defined causality setting with an automated computational process. The causality determination module uses algorithmic analysis of performance information data to substitute human expertise with machine-based automated determination.
3Reliability
If user defines causality for each system configuration change, then causality remains accurate, but management complexity increases
Solution Approach 1:
The system automatically adapts to system configuration changes by continuously collecting performance data and re-determining causality relationships. When configuration changes occur, the causality determination module autonomously analyzes the new performance information data to update causality models, eliminating the need for manual re-definition and reducing management complexity.
Solution Approach 2:
The causality model is made dynamic rather than static. The system continuously updates causality relationships based on changing system configurations by processing new performance information data, allowing the causality model to adapt automatically to dynamic environmental changes without manual intervention.
Data Source
AI summary
According to an exemplary embodiment of the present disclosure, disclosed is a computer program stored in a computer readable storage medium including encoded commands. When the computer program is executed by one or more processors of the computer system, the computer program allows the one or more processors to perform operations for generating a potential event related to an abnormal situation of an IT system. The operations include: an operation of collecting performance information data obtained by measuring values of performance indicators of a host which is a monitoring target in the IT system during a predetermined period; an operation of generating a first window with a predetermined size to be applied to the performance information data; an operation of determining a first statistical representative value of the performance information data included in the first window; an operation of generating a second window to be applied to the performance information data in which the second window has the same size as the first window and is spaced apart from the first window with a predetermined interval; an operation of determining a second statistical representative value of the performance information data included in the second window; and an operation of determining a potential event related with an abnormal situation in the IT system, at least partially based on the first statistical representative value and the second statistical representative value.


