Local Outlier Factor Hyperparameter Tuning for Anomaly Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The predictive performance of Local Outlier Factor (LOF) anomaly detection methods is hindered by the challenge of determining optimal hyperparameter values, specifically neighborhood size and contamination, which vary significantly based on the data in the training dataset, leading to suboptimal results in identifying rare or unseen anomalies.
Innovation Solution
A method is implemented to determine tuned hyperparameter values for LOF by iteratively computing LOF scores for various neighborhood size and contamination values, calculating mean and variance for outlier and inlier sets, and selecting values that maximize the difference between them, thereby improving the predictive performance of LOF models.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If default or rule-of-thumb hyperparameter values are used for LOF, then the method is simple to implement, but the predictive performance deteriorates due to variations in optimal values across different datasets
Solution Approach 1:
The patent applies parameter changes by systematically varying the neighborhood size hyperparameter across multiple candidate values and selecting the optimal value based on performance metrics. This resolves the contradiction by transforming a static default parameter approach into a dynamic optimization process that adapts parameters to specific datasets, thereby improving predictive performance while maintaining implementation simplicity through automated selection.
Solution Approach 2:
The patent implements self-service by enabling the LOF method to automatically determine its own optimal hyperparameters through internal validation mechanisms. The system performs self-tuning by evaluating different neighborhood size values and contamination ratios using cross-validation or performance metrics on validation data, allowing the method to adapt to its specific dataset without external intervention, thus resolving the contradiction between ease of implementation and predictive performance.
2Reliability
If hyperparameter tuning is performed by testing multiple values, then the predictive performance improves, but the computational time and complexity increase
Solution Approach 1:
The patent applies preliminary action by performing hyperparameter tuning during the model training phase using validation datasets before deploying the model for actual anomaly detection. By pre-determining optimal neighborhood size and contamination values through systematic evaluation on validation data, the method captures performance improvements while limiting time loss to the initial tuning phase, rather than during operational detection.
Solution Approach 2:
The patent implements partial action by testing a limited set of candidate neighborhood size values rather than exhaustively searching all possible values. The method evaluates a practical range of k values (e.g., from 5 to 50 in increments) and selects the optimal one, achieving sufficient performance improvement without the excessive computational cost of exhaustive search, thus balancing predictive performance with computational efficiency.
3Adaptability or versatility
If the neighborhood size is increased to capture more data points, then the detection of rare anomalies improves, but the precision in identifying local outliers deteriorates
Solution Approach 1:
The patent applies dynamics by making the neighborhood size a dynamic parameter that is optimized based on the specific characteristics of the dataset and the nature of anomalies to be detected. Rather than using a fixed or manually set neighborhood size, the system dynamically determines the optimal value through evaluation on validation data, allowing the method to adapt the neighborhood scale to balance detection capability and precision for different application scenarios.
Data Source
AI summary
A computing device determines hyperparameter values for outlier detection. An LOF score is computed for observation vectors using a neighborhood size value. Outlier observation vectors are selected from the observation vectors. Outlier mean and outlier variance values are computed of the LOF scores of the outlier observation vectors. Inlier observation vectors are selected from the observation vectors that have highest computed LOF scores of the observation vectors that are not included in the outlier observation vectors. Inlier mean and inlier variance values are computed of the LOF scores of the inlier observation vectors. A difference value is computed using the outlier mean and variance values and the inlier mean and variance values. The process is repeated with each neighborhood size value of a plurality of neighborhood size values. A tuned neighborhood size value is selected as the neighborhood size value associated with an extremum value of the difference value.


