Automated Database Schema Annotation via Column Similarity

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Large enterprises face inefficiencies in annotating target columns in relational databases due to the time-consuming process of data discovery, where users must manually search through thousands of tables, making it economically and time-wise challenging to employ data stewards for manual annotation.

Innovation Solution

Automated annotation techniques that utilize tabular data from sources like spreadsheets, documents, and databases to identify and annotate target columns by discovering sources, extracting tables, indexing, and calculating similarity scores between column values and names to rank and select appropriate annotations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If data stewards manually annotate target columns in relational databases, then annotation accuracy can be maintained, but the process becomes economically inefficient and time-consuming for enterprises with thousands of databases

Engineering Contradiction:
Improveannotation accuracyVSAvoidannotation time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system enables automated self-annotation of target columns by using source column names and values to automatically annotate target columns, eliminating the need for manual data steward intervention while maintaining annotation quality through similarity-based matching algorithms

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system copies annotation patterns from source columns to target columns by identifying similarities in column names and values, allowing automated replication of accurate annotations across multiple databases without manual repetition

Inventive Principle:
Principle #26Copying

2Reliability

If data stewards are employed to annotate target columns, then annotation quality can be ensured, but the cost becomes prohibitively high for enterprises with thousands of target databases

Engineering Contradiction:
Improveannotation qualityVSAvoidannotation cost
Core Design Contradiction:
ReliabilityVSLoss of energy

Solution Approach 1:

The system performs automated annotation without requiring human data stewards, using algorithmic similarity matching between source and target columns to ensure consistent quality while eliminating the high costs associated with employing specialized personnel for annotation tasks

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system replaces the mechanical process of manual annotation by data stewards with an automated computational system that uses similarity algorithms to perform annotation, thereby eliminating labor costs while maintaining annotation reliability

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Ease of operation

If users manually search through relational databases to find relevant tables, then data discovery can be performed, but the process consumes a vast majority of user time

Engineering Contradiction:
Improvedata discovery capabilityVSAvoiddata discovery time
Core Design Contradiction:
Ease of operationVSLoss of time

Solution Approach 1:

The system automatically copies and matches column information from source databases to target databases using similarity algorithms, eliminating the need for users to manually search through databases while maintaining the ability to discover relevant data

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system replaces manual user searching with an automated information retrieval system that uses similarity matching to quickly identify and annotate relevant columns, transforming the time-consuming manual search process into an automated computational task

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS10452661B2Automated database schema annotation
Publication Date: 2019.10.22 MICROSOFT TECHNOLOGY LICENSING LLC
  • US10452661B2 patent drawing
  • US10452661B2 patent drawing
  • US10452661B2 patent drawing

AI summary

Techniques and constructs that improve annotating target columns of a target database by performing automated annotation of the target columns using sources. The techniques include calculating a similarity score between a target column and columns extracted from a table that is included in a source. The similarity score is calculated based at least in part on a similarity between a value in the target column of the target database and a column value of the extracted column from the table and on a similarity between an identity of the target column of the target database and column identities of the extracted columns from the table. In some examples, the techniques calculate similarity scores for one or more extracted columns and annotate the target column based on the similarity scores.