Semantic Table Embeddings for Heterogeneous Data Retrieval

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing data unification methods, such as integrated schemas and ontologies, are inflexible and inefficient for finding compatible data across disparate data stores due to privacy and performance considerations, requiring anticipation of user needs and lacking flexibility.

Innovation Solution

A process using a trained document embedding model generates table embeddings and query embeddings, identifying responsive tables with similarity scores above a threshold, and computes compatibility and uniqueness scores to provide relevant data across heterogeneous datasets.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If integrated schemas and ontologies are used for data unification, then data compatibility across disparate data stores is improved, but system flexibility and adaptability deteriorate due to rigid pre-defined structures

Engineering Contradiction:
Improveflexibility in finding compatible dataVSAvoidcomplexity of integrated schemas and ontologies
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent replaces rigid mechanical schema-based data unification with embedding-based semantic representation. Instead of relying on pre-defined integrated schemas, the system uses document embeddings to capture semantic meaning, enabling flexible data matching across heterogeneous sources without rigid structural constraints.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent changes the fundamental parameter of data representation from structured schema-based formats to embedding vectors. This parameter transformation allows the system to adapt to any data structure by capturing semantic essence through embeddings, thereby improving flexibility while reducing the complexity of maintaining integrated schemas.

Inventive Principle:
Principle #35Parameter changes

2Productivity

If traditional data unification methods are used, then data organization is improved, but retrieval efficiency and user need anticipation deteriorate

Engineering Contradiction:
Improvedata retrieval efficiencyVSAvoidtime spent on data retrieval
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent performs preliminary action by pre-computing and storing embeddings for all dataset descriptions before queries arrive. This allows the system to quickly match new queries against pre-indexed embeddings without needing to process and interpret data structures in real-time, significantly improving retrieval efficiency.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent substitutes traditional mechanical search methods (based on schema matching and pre-defined ontologies) with semantic embedding-based retrieval. This substitution enables the system to understand and match queries based on semantic meaning rather than rigid structural patterns, improving both efficiency and adaptability to user needs.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Adaptability or versatility

If schema-based metadata is used for data description, then data structure organization is improved, but ability to handle heterogenous data across different catalogs deteriorates

Engineering Contradiction:
Improvehandling heterogenous data from different catalogsVSAvoidcomplexity of metadata heterogeneity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent introduces embedding vectors as an intermediary representation layer between raw heterogeneous data and query requirements. This intermediary transformation converts diverse data structures from multiple catalogs into a unified semantic space, enabling consistent handling and comparison of heterogenous data without requiring complex metadata normalization.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent changes the metadata representation parameter from structured schema descriptions to embedding vectors. This parameter change eliminates the complexity of handling heterogeneous metadata formats by transforming all data descriptions into a unified vector representation that captures semantic meaning, thereby simplifying the system's ability to handle data from different catalogs.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS12411840B2Embedding based heterogenous dataset evaluation
Publication Date: 2025.09.09 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US12411840B2 patent drawing
  • US12411840B2 patent drawing
  • US12411840B2 patent drawing

AI summary

An embodiment generates, using a trained document embedding model, a plurality of table embeddings, each table embedding in the plurality of table embeddings representing a table in a dataset. An embodiment generates, using the trained document embedding model, a query embedding, the query embedding representing a natural language query regarding the dataset. An embodiment identifies a set of responsive tables, the set of responsive tables comprising a first table in the dataset, the first table in the dataset represented by a first table embedding, the first table embedding and the query embedding having a similarity score above a similarity score threshold.