Automated Column Type Detection in Tabular Data Ingestion
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current Machine Learning (ML) projects face challenges due to the lack of experienced specialists, high uncertainty leading to repetitive tasks, and increased costs, particularly in data ingestion and understanding, where knowledge of column attributes like data and semantic types is crucial but often lacking.
Innovation Solution
A fully automated system, INGEST, for detecting column attributes in tabular data using a trainable and customizable feature engineering toolset with multiclass classification machine learning models, capable of handling imbalanced data and encoded/encrypted data, which extracts feature vectors from column headers and bodies to classify data and semantic types.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual data annotation and expert analysis are used to detect column attributes, then measurement precision is improved, but productivity deteriorates
Solution Approach 1:
The system enables self-service by automatically detecting column attributes through machine learning models that analyze tabular data without requiring manual expert intervention. The automated type detection system processes data independently, eliminating the need for human specialists to annotate each column while maintaining high detection accuracy.
Solution Approach 2:
The patent replaces the mechanical system of manual expert analysis with an automated machine learning-based system. Instead of relying on human specialists to manually detect column attributes, the system uses trained ML models that automatically analyze data patterns, statistical properties, and content features to determine data and semantic types.
2Measurement precision
If specialized experts are involved in data analysis, then measurement precision is improved, but device complexity deteriorates
Solution Approach 1:
The system eliminates the need for specialized experts by implementing self-service capabilities through automated machine learning models. These models independently perform semantic type detection by analyzing data patterns, statistical features, and content characteristics, replacing the need for human domain expertise in the process.
Solution Approach 2:
The patent transforms the detection process by changing parameters from manual expert judgment to automated statistical analysis. The system uses quantitative metrics such as data distribution statistics, content-based features, and pattern recognition algorithms to determine column attributes, replacing qualitative expert assessment with objective computational parameters.
3Measurement precision
If traditional type detection methods are used, then ease of operation is maintained, but measurement precision deteriorates
Solution Approach 1:
The patent replaces traditional simple detection methods with advanced machine learning-based automated detection. The system uses trained models that analyze multiple features including statistical properties, data patterns, and content characteristics to accurately determine data and semantic types, achieving high precision through automated computational analysis rather than simple rule-based approaches.
Solution Approach 2:
The system achieves both high precision and full automation by implementing self-service capabilities where machine learning models automatically detect column attributes without human intervention. The automated system processes tabular data, extracts relevant features, and determines data types and semantic types independently, maintaining ease of operation while significantly improving detection accuracy.
Data Source
AI summary
A computer-implemented method for obtaining a datasource schema comprising column-specific data-types and/or semantic-types from received tabular data records with values arranged in rows and columns, said method including: extracting a feature vector record comprising data-type recognition features for each of one or more columns of the received input tabular records; feeding the extracted feature vector records to a pretrained type classification discriminative machine learning model; using said model for classifying each extracted feature vector record of a corresponding column of received input tabular records into an estimated data-type and/or semantic-type, respectively, of the corresponding column. It is further disclosed a computer program product, a computer system and a method for training a machine learning model for obtaining the datasource schema.


