Automated Machine Log State Identification via Data Normalization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Analyzing machine logs to determine system states and predict component failures is challenging due to the voluminous and complex nature of the data, requiring extensive expertise and manual effort.
Innovation Solution
A computer-implemented method that normalizes and clusters machine log data instances to abstract common parameters, group similar data, and determine similarity distances, allowing for automated identification of system states and prediction of future errors without requiring expert input or explicit documentation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If machine logs are manually analyzed to determine system states and predict failures, then understanding of system behavior can be achieved, but the process requires extensive expertise and manual effort which becomes prohibitive with large volumes of logs
Solution Approach 1:
The system automatically analyzes machine logs and identifies system states without requiring manual expert intervention. The log analysis system processes logs, normalizes data, clusters similar log lines, and determines system states autonomously, allowing the system to serve itself in the analysis process rather than relying on external expert analysts
Solution Approach 2:
The patent replaces manual mechanical analysis processes with automated computational methods. Instead of experts manually reading and interpreting log lines, the system uses algorithms for data normalization, clustering, and state determination that automatically process large volumes of logs, substituting human cognitive effort with machine-based computational analysis
2Loss of information
If all parameters in machine logs are retained for analysis, then complete information is available for accurate state determination, but the data volume becomes prohibitive for processing
Solution Approach 1:
The system extracts and removes irrelevant parameters from machine logs during the normalization process. By identifying and eliminating parameters that do not contribute to system state determination, the system retains only the essential information needed for accurate analysis, thereby reducing data volume while preserving critical system state information
Solution Approach 2:
The patent segments the log data processing into distinct phases: normalization to remove irrelevant parameters, clustering to group similar log lines, and state determination to identify system states. This segmentation allows the system to process data in manageable stages, reducing overall data volume at each step while maintaining analytical accuracy
3Productivity
If machine logs are processed without normalization and clustering, then processing is simpler, but the ability to identify patterns and predict failures is significantly reduced
Solution Approach 1:
The system performs preliminary normalization and clustering of log data before conducting state determination and failure prediction. By pre-processing the logs to organize and group similar entries, the system prepares the data in advance for more efficient and accurate analysis, enabling both rapid processing and precise pattern recognition
Solution Approach 2:
The patent merges similar log lines into clusters based on their characteristics and patterns. By combining log entries that represent the same or similar system states, the system reduces redundancy and enhances the visibility of meaningful patterns, thereby improving both processing efficiency and prediction accuracy
Data Source
AI summary
The state of a system is determined in which data sets are generated that include a plurality of data instances representing states of one or more components of a computer system. The data instances generated by one or more data set sources that are configured to output a data instance in response to a trigger associated with the one or more components. The data instances are normalized by the application of one or more rules. The data instances from individual data set sources are separately collated to generate groups of time-specific collated data instances. State types may be assigned to each of the collated data instance groups. Distributions of state-types across the groups may be determined and a list of infrequent state-types may be generated based on the determined distributions of state-types across the groups.


