Machine Learning Model for Canonical Data Representation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing data processing techniques, such as rule-based ETL systems, are inflexible and costly, struggling to handle complex data structures and scalability, especially when aggregating data from incompatible datasets with inconsistent metadata across multiple disparate sources.

Innovation Solution

A machine learning model architecture that generates a canonical representation of datasets, using permutative input embeddings, latent representations, and alignment vectors to standardize data, enabling the aggregation of disparate datasets and generating refined predictive insights.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If rule-based ETL systems are used to aggregate data from incompatible datasets, then data transformation can be achieved, but the systems lack adaptability to changes in datasets and formatting

Engineering Contradiction:
Improveadaptability to dataset changesVSAvoidhandling of complex data structures
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent replaces the mechanical rule-based ETL system with a machine learning model that automatically learns data transformation patterns. The model uses neural networks to process unstandardized data entries and predict canonical data entities, eliminating the need for manual rule configuration and enabling automatic adaptation to changing data formats and structures.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent transforms the static rule-based parameters into dynamic learned parameters through machine learning. The model learns optimal transformation parameters from training data, allowing it to adapt to different data sources and formatting changes without manual intervention, while maintaining reliable handling of complex hierarchical data structures.

Inventive Principle:
Principle #35Parameter changes

2Extent of automation

If machine learning models are used to generate canonical representations, then adaptability and automation improve, but model training and computation resources are required

Engineering Contradiction:
Improveautomatic data standardizationVSAvoidmodel training requirements
Core Design Contradiction:
Extent of automationVSDevice complexity

Solution Approach 1:

The patent performs preliminary action by training the machine learning model in advance on comprehensive training datasets that include various data formats and structures. This pre-training enables the model to automatically handle diverse data sources without requiring complex runtime configurations or additional training, simplifying the deployment process while maintaining high automation levels.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20240411732A1Canonical transformations using machine learning language model
Publication Date: 2024.12.12 UNITEDHEALTH GROUP INC
  • US20240411732A1 patent drawing
  • US20240411732A1 patent drawing
  • US20240411732A1 patent drawing

AI summary

Various embodiments of the present disclosure provide machine learning techniques for transforming disparate, third-party datasets to canonical representations. The techniques include generating, using a machine learning prediction model, a canonical representation for an input dataset. The machine learning prediction model is previously trained using permutative input embeddings for a training dataset based on canonical data entity features, such that each permutative input embedding corresponds to a different sequence of the canonical data entity features. The permutative input embeddings are leveraged to generate a latent representation for the training dataset. The latent representation is combined with a canonical data map to generate an alignment vector, which is refined to generate an output vector for the input dataset. The machine learning prediction model is trained using a model loss generated based on a comparison of the output vector with a corresponding labeled vector.