Predicting Malfunction Recurrence via Log Anomaly Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Complexity and increased data processing in digital systems make error diagnosis and prevention difficult due to sporadic malfunctions, which can have far-reaching consequences.
Innovation Solution
A method using logged log data to identify anomalies in operating parameters by analyzing probability differences within specific time intervals, predicting imminent malfunctions through monitoring these parameters, and initiating countermeasures.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If digitalization, automation, and networking are increased to improve system functionality and data processing capability, then system performance and productivity are improved, but system complexity increases and the system becomes more prone to errors and malfunctions
Solution Approach 1:
The system performs preliminary anomaly analysis on log data to identify combinations of operating parameter values that precede malfunctions. By detecting these anomalies in advance and monitoring them continuously, the system can predict impending malfunctions before they occur, allowing preventive actions to be taken. This resolves the contradiction by adding predictive capability that increases with system complexity.
2Measurement precision
If comprehensive log data is collected and analyzed to improve malfunction detection accuracy, then measurement precision and reliability are improved, but device complexity and data processing requirements increase
Solution Approach 1:
The system extracts only the relevant information needed for anomaly detection - specifically, combinations of operating parameter values that show significant probability differences between normal and malfunction periods. Instead of analyzing all log data comprehensively, the system focuses on extracting these specific anomaly patterns, thereby improving detection accuracy while managing data processing complexity.
Solution Approach 2:
The system changes the approach from analyzing individual parameter values to analyzing combinations of parameter values. By considering joint probability distributions of multiple operating parameters simultaneously, the system can detect subtle anomalies that individual parameter analysis would miss, thereby improving detection precision without proportionally increasing complexity.
3Reliability
If anomaly analysis is performed on multiple operating parameters to improve prediction accuracy, then reliability is improved, but loss of time for data processing and analysis increases
Solution Approach 1:
The system performs anomaly analysis on a selected subset of operating parameters that are most relevant to the specific malfunction type being predicted. Instead of analyzing all possible operating parameters, the system focuses on the critical subset that shows significant probability differences, thereby maintaining high prediction accuracy while reducing the time required for data processing.
4Reliability
If continuous monitoring of operating parameters is implemented to predict impending malfunctions, then reliability and early detection capability are improved, but device complexity and computational requirements increase
Solution Approach 1:
The system uses the existing log data infrastructure and operating parameter collection mechanisms already in place, rather than requiring separate dedicated monitoring hardware or systems. By leveraging existing system resources and data collection capabilities, the anomaly detection and prediction functionality is added with minimal additional complexity, maintaining high reliability while avoiding excessive system complexity.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
The invention relates to a method for predicting a repeated occurrence of a malfunction (110) of a computer system (100, 130, 160, 198) using logged log data (122, 152, 182, 196) of the computer system (100, 130, 160, 198). The procedure comprises logging data (122, 152, 182, 196), performing an anomaly analysis for the occurrence of a malfunction (110) of the computer system (100, 130, 160, 198), identifying operating parameters (112) whose values are encompassed by one or more specific anomalies, and monitoring the values recorded for the identified operating parameters (112) in the log data records, wherein the monitoring includes predicting an impending recurrence of the malfunction (110).