Machine Learning Model for Canonical Data Representation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data processing techniques, such as rule-based ETL systems, are inflexible and costly, struggling to handle complex data structures and scalability, especially when aggregating data from incompatible datasets with inconsistent metadata across multiple disparate sources.
Innovation Solution
A machine learning model architecture that generates a canonical representation of datasets, using permutative input embeddings, latent representations, and alignment vectors to standardize data, enabling the aggregation of disparate datasets and generating refined predictive insights.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If rule-based ETL systems are used to aggregate data from incompatible datasets, then data transformation can be achieved, but the systems lack adaptability to changes in datasets and formatting
Solution Approach 1:
The patent replaces the mechanical rule-based ETL system with a machine learning model that automatically learns data transformation patterns. The model uses neural networks to process unstandardized data entries and predict canonical data entities, eliminating the need for manual rule configuration and enabling automatic adaptation to changing data formats and structures.
Solution Approach 2:
The patent transforms the static rule-based parameters into dynamic learned parameters through machine learning. The model learns optimal transformation parameters from training data, allowing it to adapt to different data sources and formatting changes without manual intervention, while maintaining reliable handling of complex hierarchical data structures.
2Extent of automation
If machine learning models are used to generate canonical representations, then adaptability and automation improve, but model training and computation resources are required
Solution Approach 1:
The patent performs preliminary action by training the machine learning model in advance on comprehensive training datasets that include various data formats and structures. This pre-training enables the model to automatically handle diverse data sources without requiring complex runtime configurations or additional training, simplifying the deployment process while maintaining high automation levels.
Data Source
AI summary
Various embodiments of the present disclosure provide machine learning techniques for transforming disparate, third-party datasets to canonical representations. The techniques include generating, using a machine learning prediction model, a canonical representation for an input dataset. The machine learning prediction model is previously trained using permutative input embeddings for a training dataset based on canonical data entity features, such that each permutative input embedding corresponds to a different sequence of the canonical data entity features. The permutative input embeddings are leveraged to generate a latent representation for the training dataset. The latent representation is combined with a canonical data map to generate an alignment vector, which is refined to generate an output vector for the input dataset. The machine learning prediction model is trained using a model loss generated based on a comparison of the output vector with a corresponding labeled vector.


