Automatic Matching Algorithm Generation for Master Data
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional master data management (MDM) systems face inefficiencies in data aggregation and centralization due to inconsistencies and multiple versions of data, requiring costly and resource-intensive processes for data synchronization and deduplication.
Innovation Solution
A method utilizing the Jaccard coefficient to determine record types and assign data records to indexing groups and comparison groups based on attribute completeness and similarity, automating the generation of matching algorithms for MDM environments, reducing the need for manual configuration and domain expertise.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Extent of automation
If conventional matching algorithms are used with manual configuration, then data matching accuracy can be maintained, but the complexity of implementation and maintenance increases significantly
Solution Approach 1:
The system enables self-service by automatically generating matching algorithms using the Jaccard coefficient without requiring manual configuration. The algorithm autonomously processes data records, determines record types, and performs deduplication tasks that previously required expert intervention and manual setup.
Solution Approach 2:
The patent replaces the mechanical/manual system of configuring matching algorithms with an automated computational system. The Jaccard coefficient-based algorithm automatically performs calculations and comparisons that previously required manual intervention, transforming a manual process into an automated computational one.
2Productivity
If manual data synchronization processes are used, then data quality can be monitored, but the time and resources required for data aggregation and centralization increase
Solution Approach 1:
The system performs preliminary action by pre-calculating and storing the Jaccard coefficient for each attribute pair before actual matching occurs. This pre-computation allows for faster real-time matching operations, as the similarity metrics are already available when data records need to be compared and deduplicated.
Solution Approach 2:
The patent replaces manual data synchronization processes with an automated algorithmic system. The matching algorithm automatically compares data records, determines duplicates, and performs consolidation without manual intervention, significantly reducing the time and resources required for data aggregation and centralization.
3Measurement precision
If comprehensive data comparison is performed across all attributes, then matching accuracy improves, but the computational cost and processing time increase
Solution Approach 1:
The patent applies segmentation by dividing the data comparison process into discrete attribute-level comparisons. Each attribute is evaluated independently using the Jaccard coefficient, allowing the system to process attributes separately and efficiently rather than comparing entire data records as monolithic units.
Solution Approach 2:
The system uses partial action by calculating the Jaccard coefficient for individual attributes rather than performing exhaustive comparisons of all possible attribute combinations. This selective approach to comparison maintains matching accuracy while reducing computational overhead by focusing on the most relevant attributes.
Data Source
AI summary
A method for receiving an additional dataset including a plurality of additional data records; determining a record type using classifiers and an internal domain knowledge corpus; dividing the plurality of additional data records into a plurality of indexing groups; assigning the given additional data record to a match set based on completeness and similarity of natures of attributes of the given additional data record; and assigning the given additional data record to and a comparison group based on completeness and similarity of natures of attributes of the given additional data record.


