Schema Generation for Data Ingestion via String Similarity
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data consolidation methods face challenges in transitioning diverse domain data models to a common data platform without loss of fidelity, leading to operational inefficiencies due to inconsistent or inaccessible data.
Innovation Solution
A method that calculates string similarities between metadata and data classes to determine data types and relationships, automatically generating a source schema for ingesting data into a common data platform, preserving data integrity and accessibility.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If data is consolidated from diverse domain data models to a common data platform, then data accessibility and operational efficiency are improved, but data fidelity and consistency may be lost
Solution Approach 1:
The system dynamically adjusts schema parameters including datatype mappings, measurement units, and relationship definitions based on string similarity calculations between source and target data models. This allows the schema to adapt to different data sources while maintaining consistency with the common data platform requirements, thereby preserving data fidelity during consolidation.
Solution Approach 2:
The patent introduces an intermediary schema generation process that acts as a mediator between diverse source data models and the common data platform. This intermediary schema, generated through automated string similarity matching and relationship assignment, ensures accurate data mapping and transformation, preventing data fidelity loss during the consolidation process.
2Productivity
If automated schema generation is implemented, then data ingestion efficiency is improved, but schema accuracy and data class determination precision may deteriorate
Solution Approach 1:
The system employs feedback mechanisms where string similarity calculation results are used to iteratively refine data class determination and attribute mapping. The calculated similarities provide feedback that guides the automated schema generation process, ensuring that data classes and attributes are accurately identified while maintaining high ingestion efficiency through automation.
Solution Approach 2:
The patent replaces manual schema creation and data class determination with automated computational methods. String similarity algorithms and automated relationship assignment replace manual analysis, achieving both high efficiency and acceptable accuracy by using computational power to perform tasks that were previously done manually with human expertise.
3Measurement precision
If string similarity calculations are performed between metadata and data classes, then data class determination accuracy is improved, but processing time and computational complexity increase
Solution Approach 1:
The patent segments the data ingestion process into distinct phases: metadata extraction, string similarity calculation, data class determination, relationship assignment, and schema generation. By segmenting the process, string similarity calculations are performed only on relevant metadata portions rather than entire datasets, reducing overall processing time while maintaining determination accuracy.
Solution Approach 2:
The system performs preliminary actions by pre-processing metadata and pre-calculating string similarities during the schema generation phase before actual data ingestion. This preliminary computation allows for faster data loading and processing later, as the classification and mapping decisions are already made, reducing the time loss during critical data ingestion operations.
Data Source
AI summary
A dataset is received from a data source. A first plurality of string similarities between metadata of the dataset with a plurality of attributes of a plurality of data classes in a target schema are calculated to determine a data class. A set of relationships are assigned to the data class based on relationships between the plurality of data classes in the target schema. A second plurality of string similarities between a plurality of attributes of the dataset and a plurality of attributes of the data class are calculated. Datatypes and measurement units are assigned to the plurality of attributes of the dataset according to the second plurality of string similarities. A source schema is generated based on the data class, the set of relationships, the plurality of attributes of the data class and the measurement units.


