Automated Data Harmonization via N-Gram and TF-IDF Mapping

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The process of harmonizing variable names across different data sets is often tedious, error-prone, and time-consuming, especially when standards allow for individual interpretation, leading to variations in naming conventions.

Innovation Solution

An automated system that maps raw data sources to a standard by analyzing variable names using n-gram analysis and term-frequency/inverse document frequency (TF-IDF) to determine the best matching model, generating a mapping rule that can be applied to new data sets, and optionally presenting user interfaces for uncertainty resolution.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If automated mapping is implemented, then productivity and accuracy are improved, but device complexity increases

Engineering Contradiction:
Improveharmonization speedVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent introduces an automated mapping system that acts as an intermediary between raw data sources and standardized data formats. This system uses n-gram analysis and TF-IDF scoring as intermediary processing layers to automatically generate mapping rules, eliminating the need for manual harmonization while managing complexity through algorithmic mediation rather than direct human intervention.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent replaces the mechanical process of manual variable name harmonization with an automated computational system. The mechanical effort of manually comparing and mapping variable names is substituted by an automated system that uses n-gram analysis, TF-IDF scoring, and machine learning models to perform the same function, thereby improving productivity while containing complexity through software automation.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Adaptability or versatility

If flexible naming conventions are allowed, then adaptability is improved, but manufacturing precision deteriorates

Engineering Contradiction:
Improvenaming flexibilityVSAvoidnaming consistency
Core Design Contradiction:
Adaptability or versatilityVSManufacturing precision

Solution Approach 1:

The patent changes the parameter of naming conventions from rigid fixed rules to flexible probabilistic mappings. By using n-gram analysis and TF-IDF scoring, the system allows multiple valid naming conventions while maintaining precision through statistical confidence thresholds. This enables adaptability to different naming styles while ensuring consistent mapping through quantified confidence levels in the automated mapping process.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent introduces dynamic mapping rules that can adapt to different data sources while maintaining overall consistency. The mapping system dynamically adjusts its behavior based on the specific characteristics of each data source, using learned patterns from training data to handle variations in naming conventions. This dynamic approach allows flexibility in accommodating different naming styles while maintaining precision through adaptive rule generation.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS10095716B1Methods, mediums, and systems for data harmonization and data harmonization and data mapping in specified domains
Publication Date: 2018.10.09 SAS INSTITUTE INC
  • US10095716B1 patent drawing
  • US10095716B1 patent drawing
  • US10095716B1 patent drawing

AI summary

The techniques described herein automatically and programmatically harmonize data, and map variable names from a dataset to standards of domains for data in the dataset. Each variable may be stored in a table which holds related groups of variables. The variables may be named by defining mappings, each mapping including two mapping rules. A first mapping rule maps a domain of the standard to the table, while a second mapping rule maps a variable within the table to a variable within the domain. When a mapping rule exists that provides an exact match between a variable name and a standard, an auto-mapping feature may be applied that automatically maps the variable name to the standard. If no exact match exists, then an analysis is performed to determine the most likely mapping candidate.