Data Quality Rule Generation for Duplicate Dataset Management
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Large organizations face challenges with duplicate datasets leading to increased infrastructure costs and inaccurate data analysis due to duplicate data storage and skewed data processing, necessitating improved database management systems.
Innovation Solution
A computing system that analyzes datasets for data quality characteristics, identifies patterns, and generates data quality rule recommendations, allowing users to implement these rules to improve data management and reduce duplication.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If duplicate datasets are stored to various databases and servers, then data availability and access flexibility are improved, but infrastructure costs and resource utilization deteriorate
Solution Approach 1:
The patent consolidates multiple duplicate datasets into a single centralized dataset, merging redundant storage locations into one unified data repository. This eliminates the need to maintain separate copies across multiple databases and servers, thereby reducing infrastructure costs while preserving data access capabilities through centralized management.
Solution Approach 2:
The centralized dataset serves as a universal data source that can be accessed by multiple teams and projects simultaneously. Instead of requiring separate duplicate datasets for different use cases, the single centralized dataset provides universal access points for various data processing needs, eliminating redundant infrastructure while maintaining adaptability.
2Reliability
If duplicate datasets are stored across multiple locations, then data redundancy and backup capability are improved, but data analysis accuracy deteriorates due to skewed analysis
Solution Approach 1:
The patent merges all duplicate datasets into a single centralized dataset, ensuring that data analysis operates on a unified, non-redundant data source. This eliminates the skewing effect caused by analyzing duplicate or inconsistent copies, thereby improving measurement precision while maintaining reliability through centralized data governance and backup mechanisms.
3Reliability
If manual data quality rule creation is performed, then data quality control is improved, but time consumption and operational efficiency deteriorate
Solution Approach 1:
The system automatically generates data quality rules by analyzing the centralized dataset and identifying patterns, anomalies, and quality issues. Instead of requiring manual creation of quality rules, the system performs self-service by autonomously detecting data quality characteristics and formulating appropriate quality rules, thereby maintaining reliable data quality control while eliminating time-consuming manual operations.
Solution Approach 2:
The system continuously analyzes the centralized dataset, provides feedback on detected data quality characteristics and patterns, and automatically adjusts quality rules based on this feedback. This closed-loop feedback mechanism ensures robust data quality control while minimizing manual intervention time, as the system learns and adapts from its own analysis results.
Data Source
AI summary
Systems and methods access, from one or more data storage locations, a dataset; perform data analysis on the dataset to detect one or more data quality characteristics each corresponding to at least one data quality dimension including timeliness, uniqueness, accuracy, completeness, validity, or consistency; evaluate the one or more data quality characteristics present in the dataset to identify one or more common patterns; and generate one or more data quality rule recommendations based on the identified one or more common patterns.


