Automated Data Quality Rule Generation via Unusual Value Combination Search
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Manual rule generation and maintenance in rule-based systems are inefficient, leading to inaccuracy and inconsistency in data quality monitoring, especially in large-scale data environments where human expertise is insufficient to handle the volume and complexity of data.
Innovation Solution
A method and apparatus for automatically identifying unusual combinations of values in data by pre-processing and searching through unique value combinations using evaluation metrics, converting these combinations into logic language rules, and employing pruning to reduce processing time and complexity.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual rule generation and maintenance is used, then domain expertise and specialized skills are utilized, but the process requires long lead times and becomes difficult to maintain consistently
Solution Approach 1:
The system performs automatic rule generation and maintenance through self-service mechanisms. The rule engine automatically discovers data quality rules, evaluates them, and maintains them without requiring continuous manual intervention, thereby reducing lead times while maintaining accuracy through automated evaluation metrics
Solution Approach 2:
The patent replaces the manual mechanical process of rule creation with an automated computational system. The rule engine uses algorithms to automatically generate, evaluate, and maintain data quality rules, substituting human manual work with automated mechanical processes that operate faster and more consistently
2Measurement precision
If manual rule creation is used, then human domain expertise is applied, but the rules become inconsistent and accuracy decreases over time
Solution Approach 1:
The system implements feedback mechanisms where the rule engine continuously evaluates generated rules against evaluation metrics and data quality standards. This feedback loop ensures rules maintain high accuracy and consistency by automatically identifying and correcting inconsistencies, preventing the degradation that occurs with manual rule maintenance
Solution Approach 2:
The patent employs parameter changes by using automated evaluation metrics with adjustable parameters to assess rule quality. The system dynamically adjusts rule parameters based on automated evaluation, ensuring consistent accuracy without the variability introduced by different human experts creating rules at different times
3Measurement precision
If exhaustive search of all value combinations is performed, then all unusual combinations are identified with high accuracy, but processing time and computational complexity increase significantly
Solution Approach 1:
The patent applies segmentation by dividing the exhaustive search space into manageable segments using pruning techniques. The rule engine segments the evaluation process by identifying and eliminating unlikely combinations early, allowing thorough analysis of promising candidates while avoiding unnecessary computation on improbable cases, thus maintaining accuracy without proportional increases in processing time
4Measurement precision
If human analysts create and maintain compliance rules, then domain knowledge is applied, but expensive consultants are required and the process becomes costly
Solution Approach 1:
The system replaces expensive human consultants with self-service automated rule generation. The rule engine automatically creates and maintains compliance rules using evaluation metrics, eliminating the need for costly external expertise while maintaining high accuracy through automated domain knowledge encoding and systematic evaluation processes
Data Source
AI summary
In a method of identifying unusual combinations of values in data (1) that is arranged in rows and columns, the data is pre-processed to put it into a form (4, 5) suitable for application of a search method thereto. Using said search method, the pre-processed data is searched (8) to search the set of possible combinations of unique values from the columns to find any combinations that are unusual according to an evaluation metric.


