Data Quality Rule Recommendations for Duplicate Dataset Control
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing database management systems face challenges with duplicate datasets leading to increased infrastructure costs and inaccurate data analysis due to the storage of duplicate data across multiple databases, which skews data analysis and leads to costly and inefficient data management.
Innovation Solution
A computing system that includes a processor, communication interface, and memory device to analyze datasets for data quality characteristics, identify patterns, and generate data quality rule recommendations, allowing for the implementation of rules to manage and improve data quality.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If duplicate datasets are stored to various databases and servers, then data availability for multiple teams is improved, but infrastructure costs increase
Solution Approach 1:
The patent implements a data sharing platform that merges multiple team datasets into a unified repository, allowing all teams to access the same data through controlled interfaces. This eliminates the need for each team to maintain separate duplicate datasets, thereby reducing overall infrastructure requirements while maintaining data availability for all teams.
Solution Approach 2:
The system creates a universal data platform that serves multiple teams and purposes simultaneously. A single stored dataset can be accessed and utilized by numerous teams for different projects and analyses, making the data storage infrastructure multi-functional and reducing the need for redundant storage across multiple systems.
2Ease of operation
If duplicate datasets are stored to various databases, then data access for multiple projects is improved, but data analysis accuracy deteriorates
Solution Approach 1:
By consolidating datasets into a unified platform, the system ensures that all teams analyze from the same source data rather than different duplicates. This merging eliminates variations between duplicate datasets that could skew analysis results, thereby improving data analysis accuracy while maintaining ease of access through the unified interface.
Solution Approach 2:
The system enforces data quality rules and validation to ensure homogeneity across all datasets stored on the platform. By standardizing data formats, quality criteria, and validation protocols, the system ensures that all teams work with consistent, high-quality data, eliminating the heterogeneity and quality variations present in separate duplicate datasets.
3Reliability
If data quality rules are implemented, then data quality is improved, but system complexity increases
Solution Approach 1:
The system implements data quality rules and validation logic at the point of data ingestion and entry, before data is fully processed or analyzed. By performing preliminary validation, cleaning, and quality checks upfront, the system ensures high data quality without requiring complex ongoing validation processes throughout the data lifecycle, thereby reducing overall system complexity.
Solution Approach 2:
The data quality system automatically detects, validates, and corrects data quality issues using predefined rules and algorithms without requiring manual intervention. The system self-monitors data quality metrics, automatically applies cleaning transformations, and flags anomalies, reducing the need for complex manual quality assurance processes and simplifying the overall system architecture.
Data Source
AI summary
Systems and methods access, from one or more data storage locations, a dataset; perform data analysis on the dataset to detect one or more data quality characteristics each corresponding to at least one data quality dimension including timeliness, uniqueness, accuracy, completeness, validity, or consistency; evaluate the one or more data quality characteristics present in the dataset to identify one or more common patterns; and generate one or more data quality rule recommendations based on the identified one or more common patterns.


