Permutation Invariant Encoding for Tabular Column Linking
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for linking tabular columns to column types in an ontology unseen during training rely on costly and time-intensive manual annotation techniques, making it impractical for businesses to efficiently manage distributed data assets across multiple custom ontologies.
Innovation Solution
The method involves encoding target tabular query columns, table headers, and target types independently to generate permutation invariant representations. This includes processing the encoded data using transformers to obtain vectors, concatenating and processing these vectors through linear and Gaussian Error Linear Unit layers to generate a final query vector, and calculating a score as a dot product between this vector and a vector representing the target types.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual annotation techniques are used to link tabular columns to column types in an ontology, then the linking accuracy can be improved, but the time consumption and cost increase significantly
Solution Approach 1:
The patent replaces manual annotation (mechanical human labor) with an automated machine learning system that uses transformers and permutation invariant encoding to link tabular columns to ontology types, thereby eliminating the time-consuming manual process while maintaining linking accuracy through algorithmic matching
Solution Approach 2:
The system enables the data linking task to serve itself by automatically processing tabular data and matching it with ontology types without requiring external manual intervention, using self-contained transformer models and encoding mechanisms to perform the entire linking workflow autonomously
2Measurement precision
If manual annotation techniques are used to link tabular columns to column types in an ontology, then the linking quality can be improved, but the cost increases significantly
Solution Approach 1:
The patent replaces expensive manual annotation processes with automated machine learning models that use permutation invariant encoding and transformer architectures to perform column-type linking, thereby reducing the financial cost associated with human expert time while maintaining high linking quality through sophisticated algorithmic matching
Solution Approach 2:
The system uses computationally efficient encoding and processing methods that can be executed quickly and at low cost using standard computing resources, replacing the need for expensive and time-intensive manual expert annotation with affordable automated processing
3Quantity of substance
If traditional methods are used to manage distributed data assets across multiple custom ontologies, then the data can be stored, but the ability to discover and visualize information is limited
Solution Approach 1:
The patent creates a universal linking mechanism that works across multiple custom ontologies and distributed data sources, enabling a single system to handle diverse data types and ontology structures, thereby improving information discovery and visualization capabilities across the entire distributed data landscape without requiring separate solutions for each ontology
Solution Approach 2:
The system introduces an intermediary layer of permutation invariant encoding and transformer-based matching that sits between the raw tabular data and the ontology types, facilitating seamless integration and discovery across multiple custom ontologies by translating diverse data formats into a unified representation that can be efficiently queried and visualized
Data Source
AI summary
An embodiment for improved linking of tabular columns to column types in an ontology unseen during training. The embodiment may for a target table, encode a target tabular query column, table headers, and target types independently to generate permutation invariant representations of tabular data associated with the target table. The embodiment may, for each of the target types, extract and further encode auxiliary information. The embodiment may process the encoded tabular data to obtain a first vector and a second vector. The embodiment may concatenate the first vector and the second vector to generate a final query vector. The embodiment may process the encoded target types through a third transformer to obtain a third vector. The embodiment may calculate a score to model interactions between the target tabular query column of the target table and the target types.


