Data Quality Rule Recommendations for Duplicate Dataset Management
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Large organizations face challenges with duplicate datasets leading to increased infrastructure costs and inaccurate data analysis due to duplicate data storage and skewed data processing, necessitating improved database management systems.
Innovation Solution
A computing system that analyzes datasets for data quality characteristics, identifies patterns, and generates data quality rule recommendations, allowing users to implement rules based on user inputs, applies these rules to new data, and determines whether to include or discard data based on confidence scores and predetermined thresholds.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If duplicate datasets are stored to various databases and servers, then data availability for multiple teams is improved, but infrastructure costs increase
Solution Approach 1:
The patent merges multiple duplicate datasets into a single centralized dataset that can be accessed by multiple teams. The system identifies duplicate datasets across different databases and servers, consolidates them into one unified dataset, and maintains data availability for all teams while eliminating redundant storage infrastructure.
2Ease of operation
If duplicate datasets are stored, then data access for multiple projects is facilitated, but data analysis accuracy deteriorates due to skewed processing
Solution Approach 1:
The system consolidates duplicate datasets into a single unified dataset, ensuring that all teams access the same data for analysis. This eliminates skewed processing results that occur when different teams analyze different versions of the same data, thereby improving data analysis accuracy while maintaining ease of access.
Solution Approach 2:
The system implements a feedback mechanism that monitors data access patterns and identifies when duplicate datasets are being created or accessed. This feedback loop allows the system to detect and resolve duplicate data issues, ensuring data consistency across all projects and maintaining analysis accuracy.
Data Source
AI summary
Systems and methods access, from one or more data storage locations, a dataset; perform data analysis on the dataset to detect one or more data quality characteristics each corresponding to at least one data quality dimension including timeliness, uniqueness, accuracy, completeness, validity, or consistency; evaluate the one or more data quality characteristics present in the dataset to identify one or more common patterns; and generate one or more data quality rule recommendations based on the identified one or more common patterns.


