Data Quality Tool Dynamic Rule Generation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data quality assessment procedures rely on manually generated, static rules that are not comprehensive and fail to account for internal and external factors, leading to inadequate data quality checks and potential false positives.
Innovation Solution
A data quality tool employing a machine learning algorithm to generate dynamic, comprehensive data quality rules based on metadata, statistical properties, and previously generated rules, which considers the influence of internal and external factors to improve data quality assessments.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual data quality rules are generated by system administrators, then the rules can be created with domain knowledge, but the process is time-consuming and the rules are not comprehensive
Solution Approach 1:
The system enables self-service by automatically generating data quality rules through machine learning algorithms that analyze historical data patterns, metadata, and statistical properties. This eliminates the need for manual rule creation by system administrators while maintaining comprehensive coverage across all data columns.
Solution Approach 2:
The patent replaces the mechanical manual process of rule generation with an automated machine learning system. The ML algorithms process data patterns, metadata, and statistical information to generate rules automatically, substituting human effort with computational intelligence.
2Adaptability or versatility
If static manually generated rules are used, then the rules are simple to implement, but they cannot adapt to internal and external factors that influence data behavior
Solution Approach 1:
The patent implements dynamic rules that automatically adapt to changing data patterns and conditions. The machine learning system continuously monitors data behavior and adjusts rules based on internal factors (data patterns, metadata changes) and external factors (business events, seasonal variations), making the system flexible and responsive.
Solution Approach 2:
The system incorporates feedback mechanisms where the machine learning model learns from historical data quality assessments and performance outcomes. This feedback loop enables the rules to improve over time by incorporating lessons from past data patterns and assessment results, enhancing adaptability to various factors.
3Reliability
If comprehensive data quality rules are manually created for all data types, then data quality assessment would be thorough, but the time and effort required is not feasible
Solution Approach 1:
The patent implements a universal machine learning-based rule generation system that handles multiple data types, formats, and domains through a single automated platform. The system analyzes metadata, statistical properties, and data patterns to generate appropriate rules for any column, eliminating the need for separate manual rule creation processes for different data types.
Solution Approach 2:
The system dynamically adjusts rule parameters based on data characteristics, statistical properties, and learned patterns. Instead of using fixed manual rules, the ML model modifies rule thresholds, conditions, and parameters automatically to match the specific characteristics of each data column, enabling comprehensive assessment across diverse data types.
Data Source
AI summary
An apparatus includes a database and a processor. The database stores a set of columns and rules assigned to each column. The rules are used to assess the quality of the data stored in the columns. The processor determines, based in part on the set of rules, the set of columns, and metadata and statistical properties of the columns, a machine learning policy adapted to generate a set of candidate rules for a given column. The processor further determines those columns of the set of columns that are similar to a subject column based on the names of the columns and the names of the tables storing the columns. The processor applies the machine learning policy to the subject column of data, rules of the similar columns, and metadata and statistical properties of the subject column to determine a set of candidate rules for the subject column.


