Outlier Detect Component for ETL Data Cleansing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Data warehouses and data marts face challenges in cleansing data due to noise and anomalies from multiple sources, which can negatively affect user interaction and decision-making processes.

Innovation Solution

The introduction of an outlier detect component in the Extract, Transform, Load (ETL) environment using a cluster mining model to differentiate between normal and outlier data, employing predictive models and artificial intelligence for accurate detection and exclusion or replacement of outliers, and providing a graphical user interface for data visualization and interaction.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If data is aggregated from multiple heterogeneous sources in data warehouses, then data volume and analytical capability are improved, but data quality deteriorates due to noise, anomalies, and outliers

Engineering Contradiction:
Improvedata volumeVSAvoiddata quality
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

The patent applies preliminary action by performing outlier detection and data cleansing operations during the ETL (Extract, Transform, Load) process before data is aggregated into the data warehouse. The outlier detect component identifies and handles anomalies in real-time during data extraction and transformation stages, preventing poor quality data from entering the warehouse while maintaining the ability to aggregate large volumes of data from multiple sources

Inventive Principle:
Principle #10Preliminary action

2Reliability

If manual data cleansing methods are used, then data quality can be improved, but productivity and time efficiency deteriorate

Engineering Contradiction:
Improvedata qualityVSAvoidtime efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent implements self-service through the automated outlier detect component that continuously monitors, detects, and handles outliers in data streams without requiring manual intervention. The system automatically identifies anomalies based on statistical models and business rules, performs cleansing operations, and maintains data quality throughout the ETL process, thereby improving both data quality and productivity simultaneously

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces manual mechanical data cleansing operations with an automated computational system. The outlier detect component uses algorithms and statistical models to automatically identify and handle anomalies, substituting human labor with computer-based detection and processing mechanisms that operate continuously and efficiently

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS7565335B2Transform for outlier detection in extract, transfer, load environment
Publication Date: 2009.07.21 MICROSOFT TECHNOLOGY LICENSING LLC
  • US7565335B2 patent drawing
  • US7565335B2 patent drawing
  • US7565335B2 patent drawing

AI summary

Systems and methods that cleanse data in Extract, Transform, Load environments (ETL), via employing an outlier detect component that is positioned in data pipeline to data warehouse(s). Such outlier detect component employs a cluster mining model to split data into normal and outlier data. Different predictive models can be employed to detect outliers in different data slices to enhance the accuracy of the predictions. In addition, a graphical user interface (GUI) enables a user to interact with cluster groups that are created and/or analyzed by the outlier detect component.