Data Deduplication Engine Configuration via Automated Profiling

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current data integration processes in master data management systems are labor-intensive and time-consuming, requiring significant manual effort and expertise to configure matching engines for deduplication, which hinders efficient data management and compliance with regulatory requirements.

Innovation Solution

A computer-implemented method and system for configuring data deduplication that analyzes source data to generate profiling statistics, classify attributes, determine data domains, and select required matching algorithms for a data matching engine, thereby automating the deduplication process and reducing manual intervention.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual configuration of matching engines is used for data deduplication, then data matching precision can be improved, but the time required and labor intensity increase significantly

Engineering Contradiction:
Improvedata matching precisionVSAvoidtime required for configuration
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs self-configuration by automatically analyzing source data, generating data profiling statistics, classifying attributes, determining data domains, and selecting matching algorithms without requiring manual intervention. This self-service approach maintains high matching precision while eliminating the time-consuming manual configuration process.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system performs preliminary analysis of source data before deduplication by generating data profiling statistics and classifying attributes in advance. This preliminary action enables the system to automatically determine data domains and select appropriate matching algorithms, avoiding the need for manual configuration while ensuring high matching precision.

Inventive Principle:
Principle #10Preliminary action

2Loss of time

If automated data analysis is implemented, then configuration time is reduced, but system complexity increases

Engineering Contradiction:
Improveconfiguration timeVSAvoidsystem complexity
Core Design Contradiction:
Loss of timeVSDevice complexity

Solution Approach 1:

The system segments the complex data analysis process into distinct modular components: data profiling statistics generation, attribute classification, data domain determination, and matching algorithm selection. Each segment handles a specific aspect of the analysis, making the overall complex system more manageable and easier to implement while achieving automated configuration.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system introduces intermediary components that facilitate automated analysis, such as data profiling statistics as intermediate representations and attribute classification as intermediary processing steps. These intermediaries simplify the relationship between input data and final matching algorithms, reducing the apparent system complexity while enabling automation.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Measurement precision

If comprehensive data analysis is performed, then data domain determination accuracy improves, but processing speed decreases

Engineering Contradiction:
Improvedata domain determination accuracyVSAvoidprocessing speed
Core Design Contradiction:
Measurement precisionVSSpeed

Solution Approach 1:

The system performs partial analysis by focusing on the most critical aspects of data profiling and attribute classification necessary for accurate data domain determination. Instead of exhaustive analysis, it selectively processes key attributes and generates essential statistics, achieving high determination accuracy while maintaining acceptable processing speed.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The system changes processing parameters dynamically based on data characteristics, adjusting the depth and scope of analysis according to the complexity and requirements of the source data. This adaptive parameter adjustment allows the system to maintain high determination accuracy while optimizing processing speed for different data scenarios.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20220374401A1Determining domain and matching algorithms for data systems
Publication Date: 2022.11.24 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US20220374401A1 patent drawing
  • US20220374401A1 patent drawing
  • US20220374401A1 patent drawing

AI summary

A computer-implemented method for configuring data deduplication is disclosed. The computer-implemented method includes receiving source data. The computer-implemented method further includes analyzing the source data, wherein analyzing the source data includes generating data profiling statistics from the source data and classifying attributes of the source data. The computer-implemented method further includes determining at least one data domain associated with the source data based, at least in part, on the data profiling statistics, the classified attributes, and ontology data. The computer-implemented method further includes determining, for the at least one data domain associated with the source data, a number of required matching algorithms for a data matching engine to execute data deduplication within the source data.