Machine Learning Data Mapping Through Cross-Column Similarity

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing predictive data analysis solutions suffer from inefficiencies and reliability issues due to the need for manual data mapping and lack of quality checks during data ingestion, leading to increased turnaround times and potential ingestion of anomalous data.

Innovation Solution

Utilizing machine learning models to perform predictive data transformation and anomaly detection by comparing column names and values, generating cross-column similarity measures, and purging anomalous values, thereby automating the data mapping process.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If manual data mapping is performed during data ingestion, then data quality can be verified, but turnaround time increases and operational efficiency decreases

Engineering Contradiction:
Improvedata qualityVSAvoidturnaround time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system performs self-service data mapping by automatically comparing incoming column names and values against the data model using machine learning, eliminating the need for manual mapping while maintaining data quality verification

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

Manual mechanical data mapping operations are replaced with automated machine learning models that analyze column names and values, substituting human effort with intelligent algorithms to reduce turnaround time while preserving quality checks

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Manufacturing precision

If manual data mapping is performed, then accurate data transformation can be achieved, but operational complexity and labor requirements increase

Engineering Contradiction:
Improvedata transformation accuracyVSAvoidoperational complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The system automatically performs data transformation mapping by itself, comparing incoming data columns against the data model without requiring manual intervention, thereby maintaining accuracy while reducing operational complexity

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

Manual data transformation operations are replaced with machine learning models that automatically compare and map columns, substituting complex manual processes with automated intelligent systems

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Reliability

If quality checks are performed during data ingestion, then anomalous data can be detected, but processing time increases

Engineering Contradiction:
Improvedata qualityVSAvoidprocessing speed
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The machine learning model performs preliminary analysis of column names and values during the ingestion process itself, detecting anomalies before full processing occurs, which maintains quality detection while minimizing additional processing time

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

Quality checks are performed continuously during the data ingestion process rather than as a separate post-processing step, maintaining uninterrupted workflow and preventing delays in processing speed

Inventive Principle:
Principle #20Continuity of useful action

Data Source

PatentUS12461900B2Machine learning techniques for enhanced data mapping
Publication Date: 2025.11.04 OPTUM INC
  • US12461900B2 patent drawing
  • US12461900B2 patent drawing
  • US12461900B2 patent drawing

AI summary

Various embodiments of the present invention provide methods, apparatus, systems, computing devices, computing entities, and/or the like for performing predictive data analysis operations. For example, certain embodiments of the present invention utilize systems, methods, and computer program products that perform predictive data transformation operations by machine learning models, where the predictive data transformation is performed based at least in part on a cross comparison of a pair of columns, accounting for the similarity of both the column names and the column values, inferred by the outputs of a machine learning model. Additionally, certain embodiments of the present invention utilize systems, methods, and computer program products that perform anomaly detection by using machine learning models that operate based at least in part on a comparison of the fixed-size representation of the column values resulting from the machine learning model.