Data Deduplication Engine Configuration via Automated Profiling
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data integration processes in master data management systems are labor-intensive and time-consuming, requiring significant manual effort and expertise to configure matching engines for deduplication, which hinders efficient data management and compliance with regulatory requirements.
Innovation Solution
A computer-implemented method and system for configuring data deduplication that analyzes source data to generate profiling statistics, classify attributes, determine data domains, and select required matching algorithms for a data matching engine, thereby automating the deduplication process and reducing manual intervention.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual configuration of matching engines is used for data deduplication, then data matching precision can be improved, but the time required and labor intensity increase significantly
Solution Approach 1:
The system performs self-configuration by automatically analyzing source data, generating data profiling statistics, classifying attributes, determining data domains, and selecting matching algorithms without requiring manual intervention. This self-service approach maintains high matching precision while eliminating the time-consuming manual configuration process.
Solution Approach 2:
The system performs preliminary analysis of source data before deduplication by generating data profiling statistics and classifying attributes in advance. This preliminary action enables the system to automatically determine data domains and select appropriate matching algorithms, avoiding the need for manual configuration while ensuring high matching precision.
2Loss of time
If automated data analysis is implemented, then configuration time is reduced, but system complexity increases
Solution Approach 1:
The system segments the complex data analysis process into distinct modular components: data profiling statistics generation, attribute classification, data domain determination, and matching algorithm selection. Each segment handles a specific aspect of the analysis, making the overall complex system more manageable and easier to implement while achieving automated configuration.
Solution Approach 2:
The system introduces intermediary components that facilitate automated analysis, such as data profiling statistics as intermediate representations and attribute classification as intermediary processing steps. These intermediaries simplify the relationship between input data and final matching algorithms, reducing the apparent system complexity while enabling automation.
3Measurement precision
If comprehensive data analysis is performed, then data domain determination accuracy improves, but processing speed decreases
Solution Approach 1:
The system performs partial analysis by focusing on the most critical aspects of data profiling and attribute classification necessary for accurate data domain determination. Instead of exhaustive analysis, it selectively processes key attributes and generates essential statistics, achieving high determination accuracy while maintaining acceptable processing speed.
Solution Approach 2:
The system changes processing parameters dynamically based on data characteristics, adjusting the depth and scope of analysis according to the complexity and requirements of the source data. This adaptive parameter adjustment allows the system to maintain high determination accuracy while optimizing processing speed for different data scenarios.
Data Source
AI summary
A computer-implemented method for configuring data deduplication is disclosed. The computer-implemented method includes receiving source data. The computer-implemented method further includes analyzing the source data, wherein analyzing the source data includes generating data profiling statistics from the source data and classifying attributes of the source data. The computer-implemented method further includes determining at least one data domain associated with the source data based, at least in part, on the data profiling statistics, the classified attributes, and ontology data. The computer-implemented method further includes determining, for the at least one data domain associated with the source data, a number of required matching algorithms for a data matching engine to execute data deduplication within the source data.


