Automated Data Harmonization via N-Gram and TF-IDF Mapping
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The process of harmonizing variable names across different data sets is often tedious, error-prone, and time-consuming, especially when standards allow for individual interpretation, leading to variations in naming conventions.
Innovation Solution
An automated system that maps raw data sources to a standard by analyzing variable names using n-gram analysis and term-frequency/inverse document frequency (TF-IDF) to determine the best matching model, generating a mapping rule that can be applied to new data sets, and optionally presenting user interfaces for uncertainty resolution.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If automated mapping is implemented, then productivity and accuracy are improved, but device complexity increases
Solution Approach 1:
The patent introduces an automated mapping system that acts as an intermediary between raw data sources and standardized data formats. This system uses n-gram analysis and TF-IDF scoring as intermediary processing layers to automatically generate mapping rules, eliminating the need for manual harmonization while managing complexity through algorithmic mediation rather than direct human intervention.
Solution Approach 2:
The patent replaces the mechanical process of manual variable name harmonization with an automated computational system. The mechanical effort of manually comparing and mapping variable names is substituted by an automated system that uses n-gram analysis, TF-IDF scoring, and machine learning models to perform the same function, thereby improving productivity while containing complexity through software automation.
2Adaptability or versatility
If flexible naming conventions are allowed, then adaptability is improved, but manufacturing precision deteriorates
Solution Approach 1:
The patent changes the parameter of naming conventions from rigid fixed rules to flexible probabilistic mappings. By using n-gram analysis and TF-IDF scoring, the system allows multiple valid naming conventions while maintaining precision through statistical confidence thresholds. This enables adaptability to different naming styles while ensuring consistent mapping through quantified confidence levels in the automated mapping process.
Solution Approach 2:
The patent introduces dynamic mapping rules that can adapt to different data sources while maintaining overall consistency. The mapping system dynamically adjusts its behavior based on the specific characteristics of each data source, using learned patterns from training data to handle variations in naming conventions. This dynamic approach allows flexibility in accommodating different naming styles while maintaining precision through adaptive rule generation.
Data Source
AI summary
The techniques described herein automatically and programmatically harmonize data, and map variable names from a dataset to standards of domains for data in the dataset. Each variable may be stored in a table which holds related groups of variables. The variables may be named by defining mappings, each mapping including two mapping rules. A first mapping rule maps a domain of the standard to the table, while a second mapping rule maps a variable within the table to a variable within the domain. When a mapping rule exists that provides an exact match between a variable name and a standard, an auto-mapping feature may be applied that automatically maps the variable name to the standard. If no exact match exists, then an analysis is performed to determine the most likely mapping candidate.


