Disambiguation Embeddings for Data Field Mapping
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing predictive data analysis solutions face inefficiencies and reliability issues in determining data table-data field relationships, particularly with unstructured data, leading to challenges in data standardization and accuracy.
Innovation Solution
A computer-implemented method and system that generates a matrix representation of a common data model, determines logical data type weights, and uses disambiguation embeddings to create input embedding vectors for predicting data field associations with candidate data tables, employing a disambiguation machine learning model to assign data fields and initiate prediction-based actions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If existing predictive data analysis solutions are used to determine data table-data field relationships, then data standardization can be performed, but computational efficiency deteriorates and training reliability is insufficient
Solution Approach 1:
The patent segments the data standardization process into multiple stages: (1) generating a matrix representation of the common data model, (2) determining logical data type weights, (3) generating disambiguation embeddings, (4) creating input embedding vectors, and (5) using a disambiguation machine learning model for final prediction. This segmentation allows each stage to process specific aspects of the data independently, improving both reliability through systematic processing and computational efficiency through targeted operations.
Solution Approach 2:
The patent performs preliminary actions by pre-processing the common data model into a matrix representation and pre-determining logical data type weights before the main prediction process. These preliminary computations prepare the data in advance, reducing the computational burden during actual prediction and improving training reliability by ensuring all necessary transformations are performed systematically beforehand.
2Manufacturing precision
If traditional methods are used for data field mapping, then data standardization is achieved, but the number of computational operations and training data entries required increases
Solution Approach 1:
The patent extracts essential information from the common data model by generating a matrix representation that captures only the critical relationships between data tables and fields. It extracts logical data type weights that summarize the importance of different data types, and generates disambiguation embeddings that capture only the necessary semantic information for accurate mapping. This extraction reduces the quantity of computational operations needed while maintaining high data standardization accuracy.
Solution Approach 2:
The patent changes the representation parameters of the data by transforming the common data model into a matrix format with specific dimensional characteristics. It transforms data field information into embedding vectors with optimized dimensions for machine learning processing. These parameter changes enable more efficient computational operations while preserving the precision needed for accurate data standardization.
3Measurement precision
If comprehensive data processing is performed, then predictive accuracy is improved, but training time and computational resources increase
Solution Approach 1:
The patent applies partial action by processing only the essential aspects of the data needed for accurate prediction. Instead of performing comprehensive processing of all data attributes, it focuses on generating the matrix representation, determining logical data type weights, and creating disambiguation embeddings for the most critical data table-field relationships. This partial processing approach achieves sufficient predictive accuracy while significantly reducing training time and computational resource requirements.
Data Source
AI summary
Various embodiments of the present disclosure provide methods, apparatus, systems, computing devices, computing entities, and/or the like for disambiguating data fields mapped to a plurality of data tables according to a common data model by generating disambiguation embeddings based on a matrix representation of the common data model and one or more logical data type weights, generating a plurality of input embedding vectors for one or more prediction inputs based on the disambiguation embeddings, generating a plurality of prediction vectors based on the plurality of input embedding vectors, and assigning one or more select data fields to respective one or more candidate data tables based on the plurality of prediction vectors.


