ML-Based Tabular Data Transformation and Cell Type Identification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing ETL mechanisms for transforming semi-structured tabular data into usable formats require multiple transformation functions, which are time-consuming and labor-intensive, necessitating an efficient and reliable method for data transformation in tabular structures.
Innovation Solution
A Machine Learning-based method that assigns scores to table cells based on orthogonal features, identifies cell and table types, and assigns unique IDs to key cells for collating analogous data from disparate tables into a central database compliant with a predefined domain schema.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If existing ETL mechanism is used to transform semi-structured tabular information, then data extraction and loading can be effectively performed, but multiple transformation functions are required which are time-consuming and require extra manual efforts
Solution Approach 1:
The system employs machine learning models that automatically analyze tabular data structures, identify cell types, and generate transformation functions without requiring manual intervention. The ML-based approach enables the system to self-configure transformation logic by learning from data patterns, thereby eliminating the time-consuming manual creation of multiple transformation functions while maintaining reliable data transformation.
Solution Approach 2:
The patent replaces the manual mechanical process of creating transformation functions with an automated machine learning system. The ML models process tabular data structures and automatically generate appropriate transformation functions, substituting the manual configuration process with an intelligent automated system that reduces time consumption while ensuring transformation reliability.
2Reliability
If existing ETL mechanism is used to transform semi-structured tabular information, then data extraction and loading can be effectively performed, but multiple transformation functions are required which require extra manual efforts
Solution Approach 1:
The system employs machine learning models that automatically analyze tabular data structures, identify cell types, and generate transformation functions without requiring manual intervention. The ML-based approach enables the system to self-configure transformation logic by learning from data patterns, thereby eliminating the need for manual creation of multiple transformation functions while maintaining reliable data transformation.
Solution Approach 2:
The patent replaces the manual mechanical process of creating transformation functions with an automated machine learning system. The ML models process tabular data structures and automatically generate appropriate transformation functions, substituting the manual configuration process with an intelligent automated system that reduces manual effort while ensuring transformation reliability.
3Productivity
If machine learning based method is used to transform tabular data, then manual effort and time are reduced, but complexity of the transformation system increases
Solution Approach 1:
The patent segments the complex data transformation task into distinct phases handled by different machine learning models: data structure analysis, cell type identification, transformation function generation, and data collation. This segmentation allows each model to specialize in a specific aspect, improving overall productivity while managing system complexity through modular architecture.
Solution Approach 2:
The patent employs universal machine learning models that can handle multiple aspects of data transformation. The same ML framework is used for analyzing different table structures, identifying various cell types, and generating diverse transformation functions, thereby reducing the need for multiple specialized systems and managing complexity through a unified approach.
Data Source
AI summary
A method for transforming data organized in tabular structure is disclosed. In some embodiments, the method includes assigning a score to each of a plurality of cells within a table based on an associated set of orthogonal features characterizing a set of data. The set of orthogonal features comprises visual features, syntactic features, and language-based features. The method further includes identifying for each of the plurality of cells, a cell type based on the assigned score. The method further includes determining a table type based on the cell type and the set of orthogonal features determined for each of the plurality of cells. The table type comprises one of a row-oriented table, a column-oriented table, or a composite table.


