Hybrid ML Anomaly Detection for Data Auditing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data auditing systems rely on sampling, which fails to capture nuances and risks, particularly in datasets with process gaps and unknown scenarios, leading to inefficient scaling due to the need for extensive human review and analysis.
Innovation Solution
A hybrid machine learning approach that includes a risk engine with a risk scorer, featuring a feature generator, ML models, and rule-based filters to assign risk scores to all data records, combining deterministic and predictive features with entity-specific rules for comprehensive risk assessment.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of time
If sampling-based auditing is used, then processing time is reduced, but measurement precision and reliability of risk identification deteriorate
Solution Approach 1:
The patent replaces manual sampling-based auditing with an automated machine learning system that processes complete datasets. The ML model automatically scores all records based on anomaly patterns, eliminating the need for human auditors to manually review sampled subsets while improving both speed and accuracy of risk identification.
Solution Approach 2:
The system enables self-service auditing where the ML model autonomously identifies and scores anomalous records without human intervention. The automated system processes entire datasets independently, generating risk scores that guide human review only when necessary, thus reducing processing time while maintaining high measurement precision.
2Measurement precision
If human auditors review all data records, then measurement precision improves, but productivity and processing speed deteriorate
Solution Approach 1:
The patent segments the auditing process into two distinct phases: (1) automated ML-based scoring of all records to identify anomalies, and (2) selective human review of only high-risk records. This segmentation allows the system to maintain high measurement precision through comprehensive analysis while improving productivity by limiting human intervention to critical cases only.
Solution Approach 2:
The system applies partial human action by having auditors review only a subset of records (those with high anomaly scores) rather than all records. The ML model performs excessive action by analyzing the complete dataset, ensuring no anomalies are missed while human resources are optimized for final verification of critical findings.
3Device complexity
If sampling methods are used, then device complexity is reduced, but adaptability to unknown scenarios and process gaps deteriorates
Solution Approach 1:
The patent replaces simple sampling mechanisms with a sophisticated ML-based system that can adapt to unknown scenarios. The model uses unsupervised learning techniques to identify novel anomaly patterns without requiring pre-defined rules, enabling the system to handle process gaps and unknown scenarios while maintaining manageable complexity through automated processing.
4Reliability
If comprehensive analysis of all records is performed, then reliability of risk identification improves, but loss of time and processing efficiency deteriorate
Solution Approach 1:
The patent segments the analysis process into automated ML scoring of all records followed by selective human review. This ensures comprehensive analysis of the entire dataset for reliable risk identification while limiting time-consuming human intervention to only high-priority cases, thus maintaining reliability without excessive time loss.
Solution Approach 2:
The ML model continuously processes all records in the dataset without interruption, ensuring comprehensive and reliable risk identification. The continuous automated analysis eliminates gaps in coverage while the system prioritizes outputs to minimize the time required for subsequent human verification steps.
Data Source
AI summary
In some aspects, the techniques described herein relate to a method including receiving, by a processor, raw data representing interactions; generating, by the processor, a feature set based on the raw data, a given feature in the feature set including at least a portion of the raw data and at least one engineered feature; generating, by the processor, a first score for the feature set using a machine learning (ML) model, the first score representing an anomaly score; generating, by the processor, one or more second scores, each score in the one or more second scores generated by performing a linear operation on one or more features in the feature set; aggregating, by the processor, the first score and the one or more second scores to generate a total score; and outputting, by the processor, the total score.


