Unsupervised ML Anomaly Detection for Hunt Data Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Managed Detection and Response (MDR) services face challenges in efficiently detecting cyberattacks due to the manual analysis of large volumes of data, which is time-consuming, error-prone, and unable to identify subtle patterns or novel attack types, leading to potential undetected breaches.

Innovation Solution

A machine learning anomaly detection system that implements a pipeline with pre-processing, dimensionality reduction, and anomaly scoring using unsupervised machine learning models, capable of identifying outlier processes or hosts without manual labeling, thereby prioritizing alerts and reducing analyst workload.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If manual data analysis is used in MDR services, then security experts can examine hunt data for compromise signals, but the process becomes extremely slow and tedious

Engineering Contradiction:
Improvedetection accuracyVSAvoidanalysis time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent replaces the manual mechanical analysis process with an automated machine learning system. The ML model automatically processes hunt data, performs feature extraction, and generates anomaly scores without human intervention, thereby eliminating the time-consuming manual analysis while maintaining or improving detection accuracy through systematic pattern recognition.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables self-service by allowing the ML model to autonomously analyze hunt data and identify anomalies without requiring continuous human oversight. The automated pipeline independently performs data preprocessing, feature engineering, anomaly scoring, and result generation, freeing security analysts from tedious manual work while preserving their ability to review critical findings.

Inventive Principle:
Principle #25Self-service

2Reliability

If manual analysis is used to detect cyberattacks, then experts can identify compromise signals, but human analysts are not always sensitive to subtle patterns and cannot easily identify novel attack types

Engineering Contradiction:
Improvedetection sensitivityVSAvoidnovel attack detection
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent employs multiple anomaly detection algorithms (Isolation Forest, One-Class SVM, Local Outlier Factor) that operate with different parameter configurations and mathematical approaches. This diversity in detection parameters enables the system to capture subtle patterns and adapt to various attack types, including novel threats, by leveraging the strengths of each algorithm in detecting different anomaly characteristics.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The ML system serves multiple detection functions simultaneously - it can identify known attack patterns through learned features while also detecting novel anomalies through unsupervised learning. The system's ability to perform both supervised and unsupervised detection makes it universally applicable to various threat types, enhancing both sensitivity and adaptability.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Productivity

If a machine learning anomaly detection system is implemented, then detection speed and accuracy are enhanced, but the system complexity increases with multiple pipeline stages

Engineering Contradiction:
Improvedetection speedVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent divides the anomaly detection system into distinct modular pipeline stages: data preprocessing, feature engineering, dimensionality reduction, anomaly scoring with multiple algorithms, and result aggregation. Each stage is independently implemented and can be selectively configured, allowing the system to achieve high detection speed through specialized processing while managing complexity through clear separation of concerns.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces intermediate components such as feature engineering layers and dimensionality reduction techniques that bridge raw data and final anomaly detection. These intermediaries simplify the overall system complexity by transforming complex raw data into structured features that are easier to process, thereby improving detection speed without requiring the final detection algorithms to handle full data complexity.

Inventive Principle:
Principle #24Intermediary (Mediator)

4Ease of manufacture

If unsupervised machine learning models are used, then manual labeling is not required and the system can learn from newly observed data, but the models need to process very large volumes of data

Engineering Contradiction:
Improvedata preparation easeVSAvoiddata volume
Core Design Contradiction:
Ease of manufactureVSQuantity of substance

Solution Approach 1:

The patent extracts essential features from large volumes of raw hunt data through feature engineering and dimensionality reduction techniques. By taking out only the most relevant features and reducing data dimensions while preserving anomaly-discriminative information, the system enables unsupervised models to learn effectively from representative subsets of data, reducing the computational burden of processing entire data volumes.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS12088600B1Machine learning system for detecting anomalies in hunt data
Publication Date: 2024.09.10 RAPID7 INC
  • US12088600B1 patent drawing
  • US12088600B1 patent drawing
  • US12088600B1 patent drawing

AI summary

An anomaly detection system is disclosed capable of reporting anomalous processes or hosts in a computer network using machine learning models trained using unsupervised training techniques. In embodiments, the system assigns observed processes to a set of process categories based on the file system path of the program executed by the process. The system extracts a feature vector for each process or host from the observation records and applies the machine learning models to the feature vectors to determine an outlier metric each process or host. The processes or hosts with the highest outlier metrics are reported as detected anomalies to be further examined by security analysts. In embodiments, the machine learnings models may be periodically retrained based on new observation records using unsupervised machine learning techniques. Accordingly, the system allows the models to learn from newly observed data without requiring the new data to be manually labeled by humans.