Automated Data Quality Detection Engine for Enterprise Systems

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Detecting data quality problems in enterprise data is challenging due to inconsistent terms and definitions across data sources, requiring specialized expertise and being time- and resource-consuming, especially when manually examining large datasets for issues like address formats and naming conventions.

Innovation Solution

An automated system for detecting potential data quality problems by processing enterprise data through cleansing, matching, and forming best records, which generates side effect data to identify non-conforming entries and provides corrective solutions, allowing users to address issues without needing extensive knowledge of various data nuances.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual examination of enterprise data is performed to detect data quality problems, then detection accuracy is improved, but time consumption and resource consumption increase significantly

Engineering Contradiction:
Improvedata quality detection accuracyVSAvoidtime consumption
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system enables self-service by allowing users to define data quality rules without requiring expert knowledge of data formats. The automated engine executes these rules and generates explanations for detected issues, making the system serve itself through automated rule execution and self-explanation capabilities.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces manual mechanical examination of data with an automated computational system. The data quality engine automatically executes detection rules, processes datasets, and generates reports without human intervention in the actual detection process, substituting human expertise with automated mechanisms.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If manual examination of enterprise data is performed to detect data quality problems, then detection accuracy is improved, but resource consumption increases significantly

Engineering Contradiction:
Improvedata quality detection accuracyVSAvoidresource consumption
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The system enables self-service by allowing users to define data quality rules without requiring expert knowledge of data formats. The automated engine executes these rules and generates explanations for detected issues, making the system serve itself through automated rule execution and self-explanation capabilities.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces manual mechanical examination of data with an automated computational system. The data quality engine automatically executes detection rules, processes datasets, and generates reports without human intervention in the actual detection process, substituting human expertise with automated mechanisms.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Productivity

If automated data processing is implemented to improve efficiency, then productivity is improved, but complexity of the system increases

Engineering Contradiction:
Improvedata processing efficiencyVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system segments data quality detection into distinct components: rule definition module, rule execution engine, data processing modules for different data types, and explanation generation. This segmentation allows each component to handle specific tasks independently, improving maintainability and reducing overall system complexity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The data quality engine provides universal functionality by handling multiple data types (addresses, phone numbers, dates) through a single unified system. The same engine can execute different detection rules for different data formats, eliminating the need for separate specialized systems for each data type.

Inventive Principle:
Principle #6Universality (Multi-functionality)

4Measurement precision

If expert knowledge of data formats is required to detect data quality problems, then detection accuracy is improved, but ease of operation deteriorates

Engineering Contradiction:
Improvedata quality detection accuracyVSAvoidease of use
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The system enables self-service by allowing users to define data quality rules without requiring expert knowledge of data formats. The automated engine executes these rules and generates explanations for detected issues, making the system serve itself through automated rule execution and self-explanation capabilities.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system introduces an intermediary layer in the form of automated rule execution and explanation generation. This intermediary translates complex data format requirements into simple user-defined rules and provides human-readable explanations for detected issues, bridging the gap between user simplicity and system complexity.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS9501504B2Automatic detection of potential data quality problems
Publication Date: 2016.11.22 SAP SE
  • US9501504B2 patent drawing
  • US9501504B2 patent drawing
  • US9501504B2 patent drawing

AI summary

Technical solutions for detection potential data quality problems are provided. In some implementations, a method includes: automatically without human intervention, identifying a subset of side effect data associated with a set of enterprise data. The side effect data include a plurality of data fields. The method further includes: selecting a first set of data quality detection rules in accordance with a first data field in the plurality of data fields; identifying one or more candidate data quality problems in the set of side effect data by comparing the set of side effect data to the first set of data quality detection rules; and responsive to identifying the one or more candidate data quality problems: causing to be displayed to a user: information representing the one or more candidate data quality problems; and one or more candidate solutions for correcting the one or more candidate data quality problems.