Disambiguation Embeddings for Data Field Mapping

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing predictive data analysis solutions face inefficiencies and reliability issues in determining data table-data field relationships, particularly with unstructured data, leading to challenges in data standardization and accuracy.

Innovation Solution

A computer-implemented method and system that generates a matrix representation of a common data model, determines logical data type weights, and uses disambiguation embeddings to create input embedding vectors for predicting data field associations with candidate data tables, employing a disambiguation machine learning model to assign data fields and initiate prediction-based actions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If existing predictive data analysis solutions are used to determine data table-data field relationships, then data standardization can be performed, but computational efficiency deteriorates and training reliability is insufficient

Engineering Contradiction:
Improvetraining reliabilityVSAvoidcomputational efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent segments the data standardization process into multiple stages: (1) generating a matrix representation of the common data model, (2) determining logical data type weights, (3) generating disambiguation embeddings, (4) creating input embedding vectors, and (5) using a disambiguation machine learning model for final prediction. This segmentation allows each stage to process specific aspects of the data independently, improving both reliability through systematic processing and computational efficiency through targeted operations.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent performs preliminary actions by pre-processing the common data model into a matrix representation and pre-determining logical data type weights before the main prediction process. These preliminary computations prepare the data in advance, reducing the computational burden during actual prediction and improving training reliability by ensuring all necessary transformations are performed systematically beforehand.

Inventive Principle:
Principle #10Preliminary action

2Manufacturing precision

If traditional methods are used for data field mapping, then data standardization is achieved, but the number of computational operations and training data entries required increases

Engineering Contradiction:
Improvedata standardization accuracyVSAvoidnumber of computational operations
Core Design Contradiction:
Manufacturing precisionVSQuantity of substance

Solution Approach 1:

The patent extracts essential information from the common data model by generating a matrix representation that captures only the critical relationships between data tables and fields. It extracts logical data type weights that summarize the importance of different data types, and generates disambiguation embeddings that capture only the necessary semantic information for accurate mapping. This extraction reduces the quantity of computational operations needed while maintaining high data standardization accuracy.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent changes the representation parameters of the data by transforming the common data model into a matrix format with specific dimensional characteristics. It transforms data field information into embedding vectors with optimized dimensions for machine learning processing. These parameter changes enable more efficient computational operations while preserving the precision needed for accurate data standardization.

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If comprehensive data processing is performed, then predictive accuracy is improved, but training time and computational resources increase

Engineering Contradiction:
Improvepredictive accuracyVSAvoidtraining time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent applies partial action by processing only the essential aspects of the data needed for accurate prediction. Instead of performing comprehensive processing of all data attributes, it focuses on generating the matrix representation, determining logical data type weights, and creating disambiguation embeddings for the most critical data table-field relationships. This partial processing approach achieves sufficient predictive accuracy while significantly reducing training time and computational resource requirements.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS20240394526A1Machine learning techniques for disambiguating unstructured data fields for mapping to data tables
Publication Date: 2024.11.28 OPTUM INC
  • US20240394526A1 patent drawing
  • US20240394526A1 patent drawing
  • US20240394526A1 patent drawing

AI summary

Various embodiments of the present disclosure provide methods, apparatus, systems, computing devices, computing entities, and/or the like for disambiguating data fields mapped to a plurality of data tables according to a common data model by generating disambiguation embeddings based on a matrix representation of the common data model and one or more logical data type weights, generating a plurality of input embedding vectors for one or more prediction inputs based on the disambiguation embeddings, generating a plurality of prediction vectors based on the plurality of input embedding vectors, and assigning one or more select data fields to respective one or more candidate data tables based on the plurality of prediction vectors.