Semantic Matching Model for Master Data De-duplication
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing master data management (MDM) systems face challenges in efficiently integrating and matching data from various sources due to complex data models, mapping, and integration processes, which lead to redundant data copies and misplaced or buried data elements.
Innovation Solution
The system employs a model-less approach using semantic label-based suspect duplicate processing and data virtualization. It generates semantic labels based on the frequency of similar values, identifies critical and non-critical data elements, auto-persists critical elements, matches them under semantic labels, establishes virtual objects for non-critical elements, and joins them with critical elements.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If data model-driven approaches are used for data mapping and integration, then data can be systematically organized, but the complexity of data mapping and integration processes increases
Solution Approach 1:
The patent extracts critical data elements (CDEs) from complex data sources and separates them from non-critical elements. By taking out only the essential matching elements and representing them as virtual objects, the system reduces mapping complexity while maintaining reliable data organization for critical information.
Solution Approach 2:
The patent segments data elements into critical data elements (CDEs) and non-critical elements based on their importance for matching and de-duplication. This segmentation allows the system to focus computational resources on organizing and matching only the essential elements, reducing overall mapping complexity while preserving data reliability.
2Productivity
If traditional data integration processes are used, then data from various sources can be combined, but redundant data copies are created
Solution Approach 1:
The patent creates virtual objects that represent critical data elements without creating physical copies of the actual data. These virtual objects are references or pointers to the source data, allowing multiple systems to access the same underlying data without duplicating it, thereby eliminating redundant data copies while maintaining integration productivity.
Solution Approach 2:
The patent introduces virtual objects as intermediaries between data sources and consuming systems. These virtual objects act as mediators that provide unified access to critical data elements from multiple sources without requiring actual data replication, thus preventing redundant copies while maintaining efficient data integration.
3Stability of the object's composition
If fixed column definitions are used for data matching, then data structure is maintained, but accuracy of data matching decreases
Solution Approach 1:
The patent dynamically identifies and profiles critical data elements from various data sources based on their actual content and importance for matching, rather than relying on fixed column definitions. This dynamic approach allows the system to adapt to different data structures while maintaining stable virtual object representations, thereby improving matching accuracy without sacrificing structural stability.
Solution Approach 2:
The patent changes the parameter of data element identification from fixed column-based definitions to dynamic content-based profiling. By analyzing actual data values, frequencies, and patterns, the system identifies true critical elements regardless of their positional parameters in source structures, improving matching accuracy while maintaining stable virtual object compositions.
4Reliability
If all data elements are processed for matching, then comprehensive coverage is achieved, but processing time increases
Solution Approach 1:
The patent extracts and identifies only the critical data elements that are essential for accurate matching and de-duplication, excluding non-critical elements from detailed processing. This extraction approach maintains matching completeness for essential elements while significantly reducing processing time by ignoring irrelevant data.
Solution Approach 2:
The patent applies partial action by processing only the critical subset of data elements necessary for matching rather than all available elements. By focusing computational resources on the essential 20% of elements that provide 80% of matching value, the system achieves reliable matching completeness with reduced processing time.
Data Source
AI summary
A computer-implemented method may include a generating a semantic label based on a frequency of similar values in a data field; and profiling a data source; identifying, based on the profiling, critical data elements (CDEs) and non-CDEs within the data source; auto-persisting the CDEs; matching the CDEs together under the semantic label; establishing virtual objects for the non-CDEs; and joining the CDEs and the virtual objects.


