Anomaly Detection via ML Classifier Models
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Manual identification of data anomalies in large, ever-changing data sets is extremely time-consuming and prone to errors, making it challenging to determine where and how data is being changed.
Innovation Solution
The method involves creating a corpus of data based on architecture or standards documents, generating classifier models, collecting information from data sources, generating feature vectors, applying these vectors to classifier models to generate analysis results, and identifying data anomalies based on these results.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual identification methods are used to detect data anomalies, then detection accuracy can be maintained through human judgment, but the process becomes extremely time-consuming and inefficient when dealing with millions of data points
Solution Approach 1:
The patent replaces manual mechanical review processes with automated machine learning models and algorithms that can process millions of data points rapidly. The system uses trained models to automatically detect anomalies, substituting human manual inspection with computational automation while maintaining or improving detection accuracy through scalable processing capabilities.
Solution Approach 2:
The system changes the parameters of data analysis by transforming raw data into feature vectors and applying multiple classification models with different thresholds and confidence levels. This allows the system to adjust detection sensitivity and processing speed dynamically, resolving the contradiction between thorough detection and time efficiency.
2Adaptability or versatility
If the data environment is constantly modified to adapt to business needs, then system adaptability improves, but determining where and how data is being changed becomes extremely challenging
Solution Approach 1:
The patent implements feedback mechanisms where the system continuously monitors data environment changes, compares them against established baselines, and automatically adjusts anomaly detection thresholds. The system provides feedback loops that track modifications to data schemas, sources, and transformations, making change detection automated rather than manual.
Solution Approach 2:
The system performs preliminary actions by establishing baseline data profiles and anomaly detection rules before changes occur. When modifications are made to the data environment, the system has pre-configured monitoring in place to immediately detect and analyze changes, rather than requiring post-change manual investigation.
3Reliability
If detection algorithms are run at multiple locations within a data pipeline to improve coverage, then anomaly detection completeness improves, but system complexity and computational overhead increase
Solution Approach 1:
The patent segments the anomaly detection system into modular components deployed at different stages of the data pipeline. Each segment handles specific detection tasks appropriate to its location, with results aggregated centrally. This segmentation allows comprehensive coverage while maintaining manageable complexity through clear separation of concerns and standardized interfaces between components.
Data Source
AI summary
A computing device may be configured to continuously, repeatedly, or recursively generate, train, improve, focus, or refine the machine learning classifier models that are used data anomalies. The computing device may create a corpus of data based on architecture or standards documents, generate classifier models based on the corpus of data, collect information from one or more data sources, generate feature vectors based on the collected information, apply the feature vectors to the classifier models to generate an analysis result, and identify a data anomaly based on the generated analysis result.


