Permutation Invariant Encoding for Tabular Column Linking

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for linking tabular columns to column types in an ontology unseen during training rely on costly and time-intensive manual annotation techniques, making it impractical for businesses to efficiently manage distributed data assets across multiple custom ontologies.

Innovation Solution

The method involves encoding target tabular query columns, table headers, and target types independently to generate permutation invariant representations. This includes processing the encoded data using transformers to obtain vectors, concatenating and processing these vectors through linear and Gaussian Error Linear Unit layers to generate a final query vector, and calculating a score as a dot product between this vector and a vector representing the target types.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual annotation techniques are used to link tabular columns to column types in an ontology, then the linking accuracy can be improved, but the time consumption and cost increase significantly

Engineering Contradiction:
Improvelinking accuracyVSAvoidannotation time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent replaces manual annotation (mechanical human labor) with an automated machine learning system that uses transformers and permutation invariant encoding to link tabular columns to ontology types, thereby eliminating the time-consuming manual process while maintaining linking accuracy through algorithmic matching

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables the data linking task to serve itself by automatically processing tabular data and matching it with ontology types without requiring external manual intervention, using self-contained transformer models and encoding mechanisms to perform the entire linking workflow autonomously

Inventive Principle:
Principle #25Self-service

2Measurement precision

If manual annotation techniques are used to link tabular columns to column types in an ontology, then the linking quality can be improved, but the cost increases significantly

Engineering Contradiction:
Improvelinking qualityVSAvoidannotation cost
Core Design Contradiction:
Measurement precisionVSLoss of energy

Solution Approach 1:

The patent replaces expensive manual annotation processes with automated machine learning models that use permutation invariant encoding and transformer architectures to perform column-type linking, thereby reducing the financial cost associated with human expert time while maintaining high linking quality through sophisticated algorithmic matching

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system uses computationally efficient encoding and processing methods that can be executed quickly and at low cost using standard computing resources, replacing the need for expensive and time-intensive manual expert annotation with affordable automated processing

Inventive Principle:
Principle #27Cheap short-living objects (Disposable)

3Quantity of substance

If traditional methods are used to manage distributed data assets across multiple custom ontologies, then the data can be stored, but the ability to discover and visualize information is limited

Engineering Contradiction:
Improvedata storage capacityVSAvoidinformation discovery capability
Core Design Contradiction:
Quantity of substanceVSEase of operation

Solution Approach 1:

The patent creates a universal linking mechanism that works across multiple custom ontologies and distributed data sources, enabling a single system to handle diverse data types and ontology structures, thereby improving information discovery and visualization capabilities across the entire distributed data landscape without requiring separate solutions for each ontology

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system introduces an intermediary layer of permutation invariant encoding and transformer-based matching that sits between the raw tabular data and the ontology types, facilitating seamless integration and discovery across multiple custom ontologies by translating diverse data formats into a unified representation that can be efficiently queried and visualized

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS12216635B2Linking tabular columns to unseen ontologies
Publication Date: 2025.02.04 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US12216635B2 patent drawing
  • US12216635B2 patent drawing
  • US12216635B2 patent drawing

AI summary

An embodiment for improved linking of tabular columns to column types in an ontology unseen during training. The embodiment may for a target table, encode a target tabular query column, table headers, and target types independently to generate permutation invariant representations of tabular data associated with the target table. The embodiment may, for each of the target types, extract and further encode auxiliary information. The embodiment may process the encoded tabular data to obtain a first vector and a second vector. The embodiment may concatenate the first vector and the second vector to generate a final query vector. The embodiment may process the encoded target types through a third transformer to obtain a third vector. The embodiment may calculate a score to model interactions between the target tabular query column of the target table and the target types.