Outlier Detect Component for ETL Data Cleansing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Data warehouses and data marts face challenges in cleansing data due to noise and anomalies from multiple sources, which can negatively affect user interaction and decision-making processes.
Innovation Solution
The introduction of an outlier detect component in the Extract, Transform, Load (ETL) environment using a cluster mining model to differentiate between normal and outlier data, employing predictive models and artificial intelligence for accurate detection and exclusion or replacement of outliers, and providing a graphical user interface for data visualization and interaction.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If data is aggregated from multiple heterogeneous sources in data warehouses, then data volume and analytical capability are improved, but data quality deteriorates due to noise, anomalies, and outliers
Solution Approach 1:
The patent applies preliminary action by performing outlier detection and data cleansing operations during the ETL (Extract, Transform, Load) process before data is aggregated into the data warehouse. The outlier detect component identifies and handles anomalies in real-time during data extraction and transformation stages, preventing poor quality data from entering the warehouse while maintaining the ability to aggregate large volumes of data from multiple sources
2Reliability
If manual data cleansing methods are used, then data quality can be improved, but productivity and time efficiency deteriorate
Solution Approach 1:
The patent implements self-service through the automated outlier detect component that continuously monitors, detects, and handles outliers in data streams without requiring manual intervention. The system automatically identifies anomalies based on statistical models and business rules, performs cleansing operations, and maintains data quality throughout the ETL process, thereby improving both data quality and productivity simultaneously
Solution Approach 2:
The patent replaces manual mechanical data cleansing operations with an automated computational system. The outlier detect component uses algorithms and statistical models to automatically identify and handle anomalies, substituting human labor with computer-based detection and processing mechanisms that operate continuously and efficiently
Data Source
AI summary
Systems and methods that cleanse data in Extract, Transform, Load environments (ETL), via employing an outlier detect component that is positioned in data pipeline to data warehouse(s). Such outlier detect component employs a cluster mining model to split data into normal and outlier data. Different predictive models can be employed to detect outliers in different data slices to enhance the accuracy of the predictions. In addition, a graphical user interface (GUI) enables a user to interact with cluster groups that are created and/or analyzed by the outlier detect component.


