Pattern-Based Data Quality Rules for Duplicate Dataset Control
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing database management systems face challenges with duplicate datasets leading to increased infrastructure costs and inaccurate data analysis due to the storage of duplicate datasets across various databases and servers, which skews data analysis and leads to costly and inefficient data management.
Innovation Solution
A computing system that includes a processor, communication interface, and memory device to analyze datasets for data quality characteristics, identify patterns, and generate data quality rule recommendations, allowing for the implementation of rules to manage and improve data quality.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If duplicate datasets are stored to various databases and servers, then data availability and access flexibility are improved, but infrastructure costs and data management complexity increase
Solution Approach 1:
The patent consolidates multiple duplicate datasets into a single centralized dataset, merging scattered data storage across various databases and servers into one unified location. This reduces data management complexity while maintaining data accessibility through centralized control and coordination.
Solution Approach 2:
The centralized dataset serves multiple teams and projects simultaneously, providing universal access to the same data source. This multi-functional approach allows different teams to access and analyze the same dataset without requiring separate duplicate copies, reducing overall system complexity.
2Ease of operation
If duplicate datasets are stored across multiple locations, then data accessibility is improved, but data accuracy and analysis reliability deteriorate due to skewing
Solution Approach 1:
By merging all duplicate datasets into a single centralized location, the patent eliminates the skewing effect that occurs when different teams analyze different versions of the same data. This ensures that all data analysis is performed on the identical dataset, improving accuracy and reliability while maintaining accessibility through centralized access protocols.
3Manufacturing precision
If data quality rules are implemented, then data accuracy and consistency are improved, but system complexity and implementation overhead increase
Solution Approach 1:
The patent implements data quality rules during the initial data collection and consolidation phase, performing preliminary validation and standardization before data is stored in the centralized dataset. This preliminary action ensures data consistency from the outset, reducing the need for complex ongoing quality control mechanisms.
Solution Approach 2:
The system incorporates automated feedback mechanisms that continuously monitor data quality metrics and enforce consistency rules. This feedback loop automatically identifies and corrects quality issues without requiring complex manual intervention, maintaining data consistency while managing system complexity through automation.
Data Source
AI summary
Systems and methods access, from one or more data storage locations, a dataset; perform data analysis on the dataset to detect one or more data quality characteristics each corresponding to at least one data quality dimension including timeliness, uniqueness, accuracy, completeness, validity, or consistency; evaluate the one or more data quality characteristics present in the dataset to identify one or more common patterns; and generate one or more data quality rule recommendations based on the identified one or more common patterns.


