Semantic Matching Model for Master Data De-duplication

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing master data management (MDM) systems face challenges in efficiently integrating and matching data from various sources due to complex data models, mapping, and integration processes, which lead to redundant data copies and misplaced or buried data elements.

Innovation Solution

The system employs a model-less approach using semantic label-based suspect duplicate processing and data virtualization. It generates semantic labels based on the frequency of similar values, identifies critical and non-critical data elements, auto-persists critical elements, matches them under semantic labels, establishes virtual objects for non-critical elements, and joins them with critical elements.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If data model-driven approaches are used for data mapping and integration, then data can be systematically organized, but the complexity of data mapping and integration processes increases

Engineering Contradiction:
Improvedata organizationVSAvoiddata mapping complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent extracts critical data elements (CDEs) from complex data sources and separates them from non-critical elements. By taking out only the essential matching elements and representing them as virtual objects, the system reduces mapping complexity while maintaining reliable data organization for critical information.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent segments data elements into critical data elements (CDEs) and non-critical elements based on their importance for matching and de-duplication. This segmentation allows the system to focus computational resources on organizing and matching only the essential elements, reducing overall mapping complexity while preserving data reliability.

Inventive Principle:
Principle #1Segmentation

2Productivity

If traditional data integration processes are used, then data from various sources can be combined, but redundant data copies are created

Engineering Contradiction:
Improvedata integration efficiencyVSAvoidredundant data copies
Core Design Contradiction:
ProductivityVSLoss of substance

Solution Approach 1:

The patent creates virtual objects that represent critical data elements without creating physical copies of the actual data. These virtual objects are references or pointers to the source data, allowing multiple systems to access the same underlying data without duplicating it, thereby eliminating redundant data copies while maintaining integration productivity.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent introduces virtual objects as intermediaries between data sources and consuming systems. These virtual objects act as mediators that provide unified access to critical data elements from multiple sources without requiring actual data replication, thus preventing redundant copies while maintaining efficient data integration.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Stability of the object's composition

If fixed column definitions are used for data matching, then data structure is maintained, but accuracy of data matching decreases

Engineering Contradiction:
Improvedata structure stabilityVSAvoiddata matching accuracy
Core Design Contradiction:
Stability of the object's compositionVSMeasurement precision

Solution Approach 1:

The patent dynamically identifies and profiles critical data elements from various data sources based on their actual content and importance for matching, rather than relying on fixed column definitions. This dynamic approach allows the system to adapt to different data structures while maintaining stable virtual object representations, thereby improving matching accuracy without sacrificing structural stability.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent changes the parameter of data element identification from fixed column-based definitions to dynamic content-based profiling. By analyzing actual data values, frequencies, and patterns, the system identifies true critical elements regardless of their positional parameters in source structures, improving matching accuracy while maintaining stable virtual object compositions.

Inventive Principle:
Principle #35Parameter changes

4Reliability

If all data elements are processed for matching, then comprehensive coverage is achieved, but processing time increases

Engineering Contradiction:
Improvematching completenessVSAvoidprocessing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent extracts and identifies only the critical data elements that are essential for accurate matching and de-duplication, excluding non-critical elements from detailed processing. This extraction approach maintains matching completeness for essential elements while significantly reducing processing time by ignoring irrelevant data.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent applies partial action by processing only the critical subset of data elements necessary for matching rather than all available elements. By focusing computational resources on the essential 20% of elements that provide 80% of matching value, the system achieves reliable matching completeness with reduced processing time.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS12314268B1Semantic matching model for data de-duplication or master data management
Publication Date: 2025.05.27 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US12314268B1 patent drawing
  • US12314268B1 patent drawing
  • US12314268B1 patent drawing

AI summary

A computer-implemented method may include a generating a semantic label based on a frequency of similar values in a data field; and profiling a data source; identifying, based on the profiling, critical data elements (CDEs) and non-CDEs within the data source; auto-persisting the CDEs; matching the CDEs together under the semantic label; establishing virtual objects for the non-CDEs; and joining the CDEs and the virtual objects.