Heterogeneous Data Mapping via Distribution Signatures

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Master data management solutions face challenges in matching and linking data from multiple sources due to variations in metadata, such as different annotations or descriptions, which hinder effective mapping and indexing.

Innovation Solution

A method and system for mapping heterogeneous data by comparing relative column positions and unique value sets, using distribution signatures to determine similarity, and generating frequency tables to normalize data, enabling accurate mapping even when metadata is not identical.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If data mapping is performed based on identical metadata annotations, then mapping accuracy is improved, but the system cannot handle heterogeneous data sources with different annotations

Engineering Contradiction:
Improvemapping accuracyVSAvoidhandling heterogeneous data sources
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The system changes the mapping approach from relying on identical metadata annotations to using distribution signatures of unique value sets. By comparing the statistical distribution of values rather than requiring identical annotations, the system achieves both high mapping accuracy and the ability to handle heterogeneous data sources with different metadata formats

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent introduces distribution signatures as an intermediary mechanism between heterogeneous data sources. Instead of directly comparing diverse metadata annotations, the system uses distribution signatures of unique value sets as a common reference frame, enabling accurate mapping across different data formats and annotation styles

Inventive Principle:
Principle #24Intermediary (Mediator)

2Ease of manufacture

If traditional mapping methods are used that rely on identical metadata, then the mapping process is simple, but it fails when metadata varies across data sources

Engineering Contradiction:
Improvesimplicity of mapping processVSAvoidmapping reliability with heterogeneous data
Core Design Contradiction:
Ease of manufactureVSReliability

Solution Approach 1:

The system transforms the mapping criterion from metadata identity to distribution signature similarity. This parameter change maintains process simplicity by using automated statistical comparison while dramatically improving reliability for heterogeneous data sources that would fail under traditional identical-metadata requirements

Inventive Principle:
Principle #35Parameter changes

3Stability of the object's composition

If exact metadata matching is required for data mapping, then data consistency is ensured, but data integration across diverse sources becomes impossible

Engineering Contradiction:
Improvedata consistencyVSAvoiddata integration capability
Core Design Contradiction:
Stability of the object's compositionVSAdaptability or versatility

Solution Approach 1:

The patent changes the consistency criterion from exact metadata matching to distribution signature equivalence. By comparing whether unique value sets have similar distributions rather than requiring identical annotations, the system ensures data consistency while enabling integration of diverse data sources with different metadata compositions

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS12050575B2Mapping of heterogeneous data as matching fields
Publication Date: 2024.07.30 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US12050575B2 patent drawing
  • US12050575B2 patent drawing
  • US12050575B2 patent drawing

AI summary

A method, a structure, and a computer system for mapping data fields. The exemplary embodiments may include, based on determining that a first data set and a second data set contain homogenous data, mapping at least one column of the first data set to at least one column of the second data set based on comparing at least one of relative column position and unique value sets. Based on determining that the first data set and the second data set contain heterogeneous data, the exemplary embodiments may include mapping the at least one column of the first data set to the at least one column of the second data set based on a difference between distribution signatures of unique value sets within each of the first data set and the second data set being less than a threshold.