Automated Database Schema Annotation via Column Similarity
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Large enterprises face inefficiencies in annotating target columns in relational databases due to the time-consuming process of data discovery, where users must manually search through thousands of tables, making it economically and time-wise challenging to employ data stewards for manual annotation.
Innovation Solution
Automated annotation techniques that utilize tabular data from sources like spreadsheets, documents, and databases to identify and annotate target columns by discovering sources, extracting tables, indexing, and calculating similarity scores between column values and names to rank and select appropriate annotations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If data stewards manually annotate target columns in relational databases, then annotation accuracy can be maintained, but the process becomes economically inefficient and time-consuming for enterprises with thousands of databases
Solution Approach 1:
The system enables automated self-annotation of target columns by using source column names and values to automatically annotate target columns, eliminating the need for manual data steward intervention while maintaining annotation quality through similarity-based matching algorithms
Solution Approach 2:
The system copies annotation patterns from source columns to target columns by identifying similarities in column names and values, allowing automated replication of accurate annotations across multiple databases without manual repetition
2Reliability
If data stewards are employed to annotate target columns, then annotation quality can be ensured, but the cost becomes prohibitively high for enterprises with thousands of target databases
Solution Approach 1:
The system performs automated annotation without requiring human data stewards, using algorithmic similarity matching between source and target columns to ensure consistent quality while eliminating the high costs associated with employing specialized personnel for annotation tasks
Solution Approach 2:
The system replaces the mechanical process of manual annotation by data stewards with an automated computational system that uses similarity algorithms to perform annotation, thereby eliminating labor costs while maintaining annotation reliability
3Ease of operation
If users manually search through relational databases to find relevant tables, then data discovery can be performed, but the process consumes a vast majority of user time
Solution Approach 1:
The system automatically copies and matches column information from source databases to target databases using similarity algorithms, eliminating the need for users to manually search through databases while maintaining the ability to discover relevant data
Solution Approach 2:
The system replaces manual user searching with an automated information retrieval system that uses similarity matching to quickly identify and annotate relevant columns, transforming the time-consuming manual search process into an automated computational task
Data Source
AI summary
Techniques and constructs that improve annotating target columns of a target database by performing automated annotation of the target columns using sources. The techniques include calculating a similarity score between a target column and columns extracted from a table that is included in a source. The similarity score is calculated based at least in part on a similarity between a value in the target column of the target database and a column value of the extracted column from the table and on a similarity between an identity of the target column of the target database and column identities of the extracted columns from the table. In some examples, the techniques calculate similarity scores for one or more extracted columns and annotate the target column based on the similarity scores.


