Automated Data Quality Rule Discovery via Conditional Dependencies
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data quality enhancement methods face challenges in efficiently identifying and managing complex data quality rules, particularly in large databases, due to exponential complexity and inability to handle noisy data, requiring significant manual effort and expertise.
Innovation Solution
A computer-implemented method generates and refines candidate conditional functional dependencies based on a dataset's ontology, applying them to data segments, refining rules to meet expectations, and selecting relevant rules to reduce complexity and handle noisy data, thereby automating the discovery of actionable data quality rules.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If automated data quality enhancement is applied to large databases, then productivity is improved, but device complexity increases due to exponential complexity in identifying complex data quality rules
Solution Approach 1:
The patent segments the database into multiple data segments and processes them in parallel using multiple computing threads. The data quality rules are applied segment-by-segment rather than to the entire database at once, dividing the complex task into manageable portions that can be processed independently and concurrently, thereby reducing overall system complexity while maintaining high productivity.
Solution Approach 2:
The patent implements dynamic rule refinement where data quality rules are initially applied broadly and then progressively refined based on results from previous segments. The system adapts and adjusts rules dynamically during processing, optimizing their application to different data segments while managing complexity through iterative improvement rather than requiring all rules to be perfectly defined upfront.
2Measurement precision
If manual SME knowledge is used to create data quality rules, then measurement precision is improved, but loss of time increases due to costly and time-consuming expert involvement
Solution Approach 1:
The patent enables the system to automatically generate and refine data quality rules without requiring manual intervention from subject matter experts. The system performs self-service by automatically analyzing data patterns, identifying anomalies, and creating rules based on observed data characteristics, thereby eliminating the time-consuming process of engaging external experts while maintaining rule accuracy through automated learning and adaptation.
Solution Approach 2:
The patent implements feedback mechanisms where the results from applying data quality rules to initial data segments are used to refine and improve subsequent rule applications. The system learns from its own performance, adjusting rules based on observed outcomes, which allows automated rule generation to achieve precision previously requiring expert knowledge while significantly reducing development time.
3Ease of operation
If traditional data profilers are used, then ease of operation is improved, but manufacturing precision deteriorates because they cannot handle noisy data or complex client-specific quality problems
Solution Approach 1:
The patent changes the operational parameters of data quality rule application by implementing noise tolerance mechanisms and adaptive rule refinement. Instead of requiring perfect, clean data as traditional profilers do, the system adjusts its parameters to handle noisy, real-world data effectively, maintaining ease of operation while significantly improving the precision and effectiveness of data quality enhancement for complex, client-specific problems.
Data Source
AI summary
Embodiments of the present invention solve the technical problem of identifying, collecting, and managing rules that improve poor quality data on enterprise initiatives ranging from data governance to business intelligence. In a specific embodiment of the present invention, a method is provided for producing data quality rules for a data set. A set of candidate conditional functional dependencies are generated comprised of candidate seeds of attributes that are within a certain degree of relatedness in the ontology of the data set. The candidate conditional functional dependencies are then applied to the data refined until they reach a quiescent state where they have not been refined even though the data they have been applied to has been stable. The resulting refined candidate conditional functional dependencies are the data enhancement rules for the data set and other related data sets. In another specific embodiment of the present invention, a computer system for the development of data quality rules is provided having a rule repository, a data quality rules discovery engine, and a user interface.


