AI Data Anomaly Detection With Self-Correcting Rule Reconfiguration
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Organizations face challenges in managing data quality due to anomalies, biases, and noise in datasets, which affect the performance and reliability of AI models, leading to flawed analytics and operational risks.
Innovation Solution
A data management platform using AI models to identify, evaluate, and correct anomalies by comparing observed patterns with reference patterns, generating reconfiguration commands to align with expected rules, and detecting out-of-distribution data to improve data quality and reliability of AI models.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If manual data cleaning methods are used, then data quality can be improved, but time and resources required increase significantly
Solution Approach 1:
The system enables self-service by having the AI model automatically detect and correct its own anomalies through the self-diagnosis module, eliminating the need for manual intervention in the data cleaning process
Solution Approach 2:
The patent replaces manual mechanical data cleaning processes with an automated AI-based system that uses machine learning models to detect and correct anomalies, substituting human labor with intelligent automation
2Reliability
If AI models are trained on datasets with anomalies, then model performance decreases, but data cleaning complexity increases
Solution Approach 1:
The system performs preliminary action by detecting and correcting anomalies in the training data before the AI model is trained, ensuring that the model receives clean data as input and preventing anomaly-induced performance degradation
Solution Approach 2:
The self-diagnosis module provides feedback by automatically identifying anomalies in the training data and generating corrections, creating a closed-loop system that continuously improves data quality without increasing operational complexity
3Measurement precision
If more data is collected to improve training accuracy, then model accuracy improves, but the proportion of irrelevant data increases
Solution Approach 1:
The system extracts and removes irrelevant data by using the anomaly detection module to identify and filter out noisy or irrelevant data points from the training dataset, allowing the model to be trained on a cleaner, more relevant subset of data
Solution Approach 2:
The system changes the parameter of data relevance by transforming the raw dataset into a cleaned version with improved quality metrics, effectively altering the composition and characteristics of the training data
Data Source
AI summary
The systems and methods disclosed herein receive a dataset including an observed set of values for a set of variables. The system can use a first set of AI models to identify a set of anomalies in the observed set of values by comparing an observed set of patterns against multiple reference patterns. The system can use a second set of AI models to evaluate the identified anomalies by comparing an observed set of association rules with an expected set of association rules. The system can use a third set of AI models to generate reconfiguration commands to remove the identified anomalies. The reconfiguration commands can be automatically executed to modify the observed association rules to align with the expected association rules.


