Embedding-Based Schema Matching for Disparate Data Sources

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing data analytics systems struggle to identify and align similar fields across disparate data sources due to differing schemas and manual methods requiring domain expertise, leading to inefficiencies in data combination and integration.

Innovation Solution

A system utilizing multi-dimensional matrices, external embeddings, transformation operations, and similarity metrics to generate embeddings sets and representation vectors, enabling automated schema matching between multiple data sources.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual methods are used to identify and align similar fields across disparate data sources, then domain expertise can be applied to ensure accuracy, but the process becomes time-consuming and requires significant human effort

Engineering Contradiction:
Improveschema matching accuracyVSAvoidtime for field alignment
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent introduces an intermediary system that acts as a mediator between disparate data sources. This system uses embedding models to transform data from different schemas into a common semantic space, enabling automatic matching without requiring manual domain expertise while maintaining high accuracy through learned representations

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent replaces the mechanical manual process of schema matching with an automated computational system. Instead of human experts manually comparing and aligning fields, the system uses machine learning models to automatically transform and match schemas, significantly reducing time while maintaining or improving accuracy

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Productivity

If basic methods are used to identify similar fields by selecting or looking at fields with similar names, then the process is simple and quick, but columns with similar names can store data with diverse meanings leading to incorrect matching

Engineering Contradiction:
Improvedata integration speedVSAvoidfield similarity identification accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent transforms the matching parameters from superficial string comparison to deep semantic representations. By changing the parameter from field name similarity to embedding space similarity, the system can quickly process data while accurately capturing the true meaning and context of each field, avoiding false matches based solely on naming conventions

Inventive Principle:
Principle #35Parameter changes

3Extent of automation

If automated embedding-based methods are used for schema matching, then manual effort and domain expertise requirements are reduced, but the system complexity increases due to multi-dimensional matrices and transformation operations

Engineering Contradiction:
Improveschema matching automationVSAvoidsystem architectural complexity
Core Design Contradiction:
Extent of automationVSDevice complexity

Solution Approach 1:

The patent segments the complex schema matching process into distinct modular components: embedding generation, transformation operations, and similarity computation. Each component handles a specific aspect of the matching process, making the overall system more manageable and maintainable while achieving high automation

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent creates a universal embedding-based framework that can handle multiple data sources and schemas through a single unified approach. The transformation operations and similarity metrics serve multiple purposes across different matching scenarios, reducing overall system complexity despite the automated nature of the process

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS12373404B2Schema matching of disparate data sources using embedding-based similarity scores
Publication Date: 2025.07.29 OPTUM SERVICES IRELAND LTD
  • US12373404B2 patent drawing
  • US12373404B2 patent drawing
  • US12373404B2 patent drawing

AI summary

Embodiments provide for schema matching between multiple disparate data sources using multi-dimensional matrices, external embeddings, transformation operations, and similarity metrics.