Automated Data Quality Detection Engine for Enterprise Systems
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Detecting data quality problems in enterprise data is challenging due to inconsistent terms and definitions across data sources, requiring specialized expertise and being time- and resource-consuming, especially when manually examining large datasets for issues like address formats and naming conventions.
Innovation Solution
An automated system for detecting potential data quality problems by processing enterprise data through cleansing, matching, and forming best records, which generates side effect data to identify non-conforming entries and provides corrective solutions, allowing users to address issues without needing extensive knowledge of various data nuances.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual examination of enterprise data is performed to detect data quality problems, then detection accuracy is improved, but time consumption and resource consumption increase significantly
Solution Approach 1:
The system enables self-service by allowing users to define data quality rules without requiring expert knowledge of data formats. The automated engine executes these rules and generates explanations for detected issues, making the system serve itself through automated rule execution and self-explanation capabilities.
Solution Approach 2:
The patent replaces manual mechanical examination of data with an automated computational system. The data quality engine automatically executes detection rules, processes datasets, and generates reports without human intervention in the actual detection process, substituting human expertise with automated mechanisms.
2Measurement precision
If manual examination of enterprise data is performed to detect data quality problems, then detection accuracy is improved, but resource consumption increases significantly
Solution Approach 1:
The system enables self-service by allowing users to define data quality rules without requiring expert knowledge of data formats. The automated engine executes these rules and generates explanations for detected issues, making the system serve itself through automated rule execution and self-explanation capabilities.
Solution Approach 2:
The patent replaces manual mechanical examination of data with an automated computational system. The data quality engine automatically executes detection rules, processes datasets, and generates reports without human intervention in the actual detection process, substituting human expertise with automated mechanisms.
3Productivity
If automated data processing is implemented to improve efficiency, then productivity is improved, but complexity of the system increases
Solution Approach 1:
The system segments data quality detection into distinct components: rule definition module, rule execution engine, data processing modules for different data types, and explanation generation. This segmentation allows each component to handle specific tasks independently, improving maintainability and reducing overall system complexity.
Solution Approach 2:
The data quality engine provides universal functionality by handling multiple data types (addresses, phone numbers, dates) through a single unified system. The same engine can execute different detection rules for different data formats, eliminating the need for separate specialized systems for each data type.
4Measurement precision
If expert knowledge of data formats is required to detect data quality problems, then detection accuracy is improved, but ease of operation deteriorates
Solution Approach 1:
The system enables self-service by allowing users to define data quality rules without requiring expert knowledge of data formats. The automated engine executes these rules and generates explanations for detected issues, making the system serve itself through automated rule execution and self-explanation capabilities.
Solution Approach 2:
The system introduces an intermediary layer in the form of automated rule execution and explanation generation. This intermediary translates complex data format requirements into simple user-defined rules and provides human-readable explanations for detected issues, bridging the gap between user simplicity and system complexity.
Data Source
AI summary
Technical solutions for detection potential data quality problems are provided. In some implementations, a method includes: automatically without human intervention, identifying a subset of side effect data associated with a set of enterprise data. The side effect data include a plurality of data fields. The method further includes: selecting a first set of data quality detection rules in accordance with a first data field in the plurality of data fields; identifying one or more candidate data quality problems in the set of side effect data by comparing the set of side effect data to the first set of data quality detection rules; and responsive to identifying the one or more candidate data quality problems: causing to be displayed to a user: information representing the one or more candidate data quality problems; and one or more candidate solutions for correcting the one or more candidate data quality problems.


