Log Anomaly Detection Using Unlabeled Data Clustering
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Modern anomaly detection systems in initial deployment situations, known as "0-day" scenarios, are inefficient due to scarce training data and high false positive rates, leading to reduced operational efficiency and increased computational wastage.
Innovation Solution
A computer-implemented method that classifies log lines, templatizes, clusters, and trains a log anomaly model using unlabeled data to identify anomalous log lines, reducing false positives and requiring less training data, while correlating anomalous events across environments to pinpoint causes and suggest remedies.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional anomaly detection systems are deployed in initial deployment situations, then they can detect anomalies, but they produce high false positive rates and require extensive training data
Solution Approach 1:
The system performs preliminary classification of log lines into erroneous and non-erroneous categories before training the anomaly detection model. This preliminary action enables the system to create initial training data from unlabeled logs, reducing the need for extensive manually labeled training data while improving detection reliability in initial deployment scenarios
Solution Approach 2:
The system uses its own operational logs to automatically generate training data through classification and templatization. By self-service, the system converts its own log data into labeled training examples, eliminating the need for external extensive training data collection and enabling effective anomaly detection from the start
2Reliability
If traditional anomaly detection systems are deployed in initial deployment situations, then they can detect anomalies, but they experience reduced operational efficiency due to high false positives
Solution Approach 1:
The system performs preliminary classification and templatization of log lines before anomaly detection. This preliminary action creates structured training data that improves model accuracy, thereby reducing false positives and enhancing operational efficiency by enabling more reliable anomaly detection from the outset
Solution Approach 2:
The system implements continuous learning where anomaly detection results and log classifications feed back into model retraining. This feedback mechanism continuously improves detection accuracy and reduces false positives over time, maintaining high operational efficiency while preserving reliable anomaly detection
3Reliability
If traditional anomaly detection systems are deployed in initial deployment situations, then they can detect anomalies, but they incur increased computational wastage
Solution Approach 1:
The system performs preliminary classification and templatization of log lines to create structured training data before anomaly detection. This preliminary action improves model effectiveness, reducing the need for extensive computational resources during anomaly detection operations and minimizing computational wastage while maintaining detection reliability
Solution Approach 2:
The system changes the state of log data from raw unlabeled format to classified and templatized structured format. This parameter change in data organization improves model training efficiency and reduces computational wastage during anomaly detection while preserving accurate anomaly detection capability
Data Source
AI summary
One or more computer processors classify each log line in a plurality of unlabeled log lines as an erroneous log line or a non-erroneous log line. The one or more computer processors templatize each classified erroneous log line and non-erroneous log line in the plurality of unlabeled log lines. The one or more computer processors cluster erroneous log templates into erroneous log template clusters and the non-erroneous log templates into non-erroneous log template clusters. The one or more computer processors eliminate the erroneous log template clusters and the non-erroneous log template clusters that exceed a frequency threshold. The one or more computer processors train a log anomaly model utilizing=remaining erroneous log template clusters and remaining non-erroneous log template clusters. The one or more computer processors identify a subsequent log line as anomalous or non-anomalous utilizing the trained log anomaly model.


