Duplicate Data Record Consolidation via ACE Enrichment
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Large data processing environments face challenges in maintaining data accuracy and consistency due to duplicate data records, which can limit data aggregation and updating processes, and often require manual handling, leading to increased processing needs and security risks.
Innovation Solution
A system that detects and updates duplicate data records by designating an Active Consolidation Entity (ACE) record, enriching it with data from other duplicate records, and overwriting the remaining records with ACE data, thereby reducing storage needs and improving data integrity.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If duplicate data records are removed from the data processing environment, then data accuracy and consistency are improved, but downstream consumer systems that relied on the removed duplicate data record will no longer have valid requests
Solution Approach 1:
The patent extracts the core data from duplicate records and consolidates it into a single canonical record, while separating out the reference relationships to downstream systems. This allows the duplicate data to be removed without breaking downstream dependencies, as the canonical record serves as the single source of truth that all downstream systems can reference.
Solution Approach 2:
The patent introduces an intermediary mechanism (the canonical record with reference tracking) that mediates between the need to remove duplicates and the need to maintain downstream system compatibility. The canonical record acts as a mediator that preserves data integrity while maintaining valid reference paths for downstream consumers.
2Reliability
If manual handling of duplicate data records is performed, then data quality can be maintained, but processing time and security risks increase
Solution Approach 1:
The patent implements self-service through automated detection and resolution of duplicate records. The system automatically identifies duplicates, selects canonical records using defined criteria, and updates references without human intervention. This eliminates manual processing while maintaining data quality through systematic, rule-based operations.
Solution Approach 2:
The patent incorporates feedback mechanisms where the system continuously monitors data quality metrics and automatically adjusts duplicate detection and resolution processes. The feedback loop ensures that automated processes maintain the same quality standards as manual handling would provide, while eliminating time losses.
3Adaptability or versatility
If duplicate data records are kept in the data processing environment, then downstream consumer systems can continue to access data, but data aggregation accuracy and completeness are limited
Solution Approach 1:
The patent merges duplicate data records into a single canonical record while preserving accessibility through reference mechanisms. All duplicate records are consolidated into one authoritative source, ensuring that data aggregation operates on unique, non-duplicated data. Downstream systems maintain access through references to the canonical record, achieving both data accuracy and accessibility.
4Quantity of substance
If data records from uncertified data sources are ingested, then data volume increases, but data quality decreases leading to more duplicate records
Solution Approach 1:
The patent applies parameter changes by implementing quality thresholds and certification criteria that filter incoming data. Records from uncertified sources undergo additional validation checks, and duplicate detection sensitivity is adjusted based on source reliability. This allows the system to ingest large volumes of data while maintaining quality standards through dynamic parameter adjustment.
Data Source
AI summary
Systems, methods, and articles of manufacture for detecting and updating duplicate data records are provided. The system may be configured to detect and retrieve duplicate data records in a data storage and generate a data duplicate reference set comprising the duplicate data records. The duplicate data records may be grouped into common data duplicate groups within the data duplicate reference set. The system may elect one of the duplicate data records from one of the data duplicate groups to be an ACE record. The ACE record may be enriched using the remaining duplicate data records. Data from each duplicate data record may then be overwritten using the ACE record data to ensure that all of the duplicate data records comprise the same data. All of the duplicate data records may be cross-linked in the data storage to ensure consistency and data integrity throughout the data duplicate group.


