Unsupervised ML Anomaly Detection for Hunt Data Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Managed Detection and Response (MDR) services face challenges in efficiently detecting cyberattacks due to the manual analysis of large volumes of data, which is time-consuming, error-prone, and unable to identify subtle patterns or novel attack types, leading to potential undetected breaches.
Innovation Solution
A machine learning anomaly detection system that implements a pipeline with pre-processing, dimensionality reduction, and anomaly scoring using unsupervised machine learning models, capable of identifying outlier processes or hosts without manual labeling, thereby prioritizing alerts and reducing analyst workload.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If manual data analysis is used in MDR services, then security experts can examine hunt data for compromise signals, but the process becomes extremely slow and tedious
Solution Approach 1:
The patent replaces the manual mechanical analysis process with an automated machine learning system. The ML model automatically processes hunt data, performs feature extraction, and generates anomaly scores without human intervention, thereby eliminating the time-consuming manual analysis while maintaining or improving detection accuracy through systematic pattern recognition.
Solution Approach 2:
The system enables self-service by allowing the ML model to autonomously analyze hunt data and identify anomalies without requiring continuous human oversight. The automated pipeline independently performs data preprocessing, feature engineering, anomaly scoring, and result generation, freeing security analysts from tedious manual work while preserving their ability to review critical findings.
2Reliability
If manual analysis is used to detect cyberattacks, then experts can identify compromise signals, but human analysts are not always sensitive to subtle patterns and cannot easily identify novel attack types
Solution Approach 1:
The patent employs multiple anomaly detection algorithms (Isolation Forest, One-Class SVM, Local Outlier Factor) that operate with different parameter configurations and mathematical approaches. This diversity in detection parameters enables the system to capture subtle patterns and adapt to various attack types, including novel threats, by leveraging the strengths of each algorithm in detecting different anomaly characteristics.
Solution Approach 2:
The ML system serves multiple detection functions simultaneously - it can identify known attack patterns through learned features while also detecting novel anomalies through unsupervised learning. The system's ability to perform both supervised and unsupervised detection makes it universally applicable to various threat types, enhancing both sensitivity and adaptability.
3Productivity
If a machine learning anomaly detection system is implemented, then detection speed and accuracy are enhanced, but the system complexity increases with multiple pipeline stages
Solution Approach 1:
The patent divides the anomaly detection system into distinct modular pipeline stages: data preprocessing, feature engineering, dimensionality reduction, anomaly scoring with multiple algorithms, and result aggregation. Each stage is independently implemented and can be selectively configured, allowing the system to achieve high detection speed through specialized processing while managing complexity through clear separation of concerns.
Solution Approach 2:
The patent introduces intermediate components such as feature engineering layers and dimensionality reduction techniques that bridge raw data and final anomaly detection. These intermediaries simplify the overall system complexity by transforming complex raw data into structured features that are easier to process, thereby improving detection speed without requiring the final detection algorithms to handle full data complexity.
4Ease of manufacture
If unsupervised machine learning models are used, then manual labeling is not required and the system can learn from newly observed data, but the models need to process very large volumes of data
Solution Approach 1:
The patent extracts essential features from large volumes of raw hunt data through feature engineering and dimensionality reduction techniques. By taking out only the most relevant features and reducing data dimensions while preserving anomaly-discriminative information, the system enables unsupervised models to learn effectively from representative subsets of data, reducing the computational burden of processing entire data volumes.
Data Source
AI summary
An anomaly detection system is disclosed capable of reporting anomalous processes or hosts in a computer network using machine learning models trained using unsupervised training techniques. In embodiments, the system assigns observed processes to a set of process categories based on the file system path of the program executed by the process. The system extracts a feature vector for each process or host from the observation records and applies the machine learning models to the feature vectors to determine an outlier metric each process or host. The processes or hosts with the highest outlier metrics are reported as detected anomalies to be further examined by security analysts. In embodiments, the machine learnings models may be periodically retrained based on new observation records using unsupervised machine learning techniques. Accordingly, the system allows the models to learn from newly observed data without requiring the new data to be manually labeled by humans.


