Automated Column Type Detection in Tabular Data Ingestion

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current Machine Learning (ML) projects face challenges due to the lack of experienced specialists, high uncertainty leading to repetitive tasks, and increased costs, particularly in data ingestion and understanding, where knowledge of column attributes like data and semantic types is crucial but often lacking.

Innovation Solution

A fully automated system, INGEST, for detecting column attributes in tabular data using a trainable and customizable feature engineering toolset with multiclass classification machine learning models, capable of handling imbalanced data and encoded/encrypted data, which extracts feature vectors from column headers and bodies to classify data and semantic types.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual data annotation and expert analysis are used to detect column attributes, then measurement precision is improved, but productivity deteriorates

Engineering Contradiction:
Improvecolumn attribute detection accuracyVSAvoiddata ingestion speed
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The system enables self-service by automatically detecting column attributes through machine learning models that analyze tabular data without requiring manual expert intervention. The automated type detection system processes data independently, eliminating the need for human specialists to annotate each column while maintaining high detection accuracy.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical system of manual expert analysis with an automated machine learning-based system. Instead of relying on human specialists to manually detect column attributes, the system uses trained ML models that automatically analyze data patterns, statistical properties, and content features to determine data and semantic types.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If specialized experts are involved in data analysis, then measurement precision is improved, but device complexity deteriorates

Engineering Contradiction:
Improvesemantic type detection accuracyVSAvoidsystem expertise requirement
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system eliminates the need for specialized experts by implementing self-service capabilities through automated machine learning models. These models independently perform semantic type detection by analyzing data patterns, statistical features, and content characteristics, replacing the need for human domain expertise in the process.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent transforms the detection process by changing parameters from manual expert judgment to automated statistical analysis. The system uses quantitative metrics such as data distribution statistics, content-based features, and pattern recognition algorithms to determine column attributes, replacing qualitative expert assessment with objective computational parameters.

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If traditional type detection methods are used, then ease of operation is maintained, but measurement precision deteriorates

Engineering Contradiction:
Improvedata type detection accuracyVSAvoidautomation level
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The patent replaces traditional simple detection methods with advanced machine learning-based automated detection. The system uses trained models that analyze multiple features including statistical properties, data patterns, and content characteristics to accurately determine data and semantic types, achieving high precision through automated computational analysis rather than simple rule-based approaches.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system achieves both high precision and full automation by implementing self-service capabilities where machine learning models automatically detect column attributes without human intervention. The automated system processes tabular data, extracts relevant features, and determines data types and semantic types independently, maintaining ease of operation while significantly improving detection accuracy.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS20230316147A1Method and system for obtaining a datasource schema comprising column-specific data-types and/or semantic-types from received tabular data records
Publication Date: 2023.10.05 FEEDZAI CONSULTADORIA E INOVACAO TECHCA SA
  • US20230316147A1 patent drawing
  • US20230316147A1 patent drawing
  • US20230316147A1 patent drawing

AI summary

A computer-implemented method for obtaining a datasource schema comprising column-specific data-types and/or semantic-types from received tabular data records with values arranged in rows and columns, said method including: extracting a feature vector record comprising data-type recognition features for each of one or more columns of the received input tabular records; feeding the extracted feature vector records to a pretrained type classification discriminative machine learning model; using said model for classifying each extracted feature vector record of a corresponding column of received input tabular records into an estimated data-type and/or semantic-type, respectively, of the corresponding column. It is further disclosed a computer program product, a computer system and a method for training a machine learning model for obtaining the datasource schema.