Automated Database Data Type Classification via Machine Learning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for data modernization face challenges in determining data types in legacy databases due to limited metadata and the inefficiency of manual labeling by subject matter experts, which complicates the process of moving data to modern databases.

Innovation Solution

A computer-implemented method using machine learning models to predict data types by generating descriptions from partial database component information, expanding acronyms and abbreviations, and utilizing trained models with labeled data to classify data types automatically.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual labeling by subject matter experts is used to determine data types, then accuracy of data type classification can be maintained, but productivity is reduced and loss of time increases

Engineering Contradiction:
Improveaccuracy of data type classificationVSAvoidproductivity
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The system enables self-service by allowing the database components themselves to provide sufficient information for data type classification through their identifiers, descriptions, and metadata. The machine learning model processes this self-provided information to automatically determine data types without requiring external expert intervention, thus resolving the contradiction between maintaining accuracy and improving productivity.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical system of manual expert labeling with an automated machine learning-based classification system. The ML model processes database component information and automatically predicts data types, substituting human expert manual work with an automated computational process that maintains accuracy while significantly improving productivity.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If manual labeling by subject matter experts is used to determine data types, then accuracy of data type classification can be maintained, but loss of time increases

Engineering Contradiction:
Improveaccuracy of data type classificationVSAvoidloss of time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system enables self-service by allowing the database components themselves to provide sufficient information for data type classification through their identifiers, descriptions, and metadata. The machine learning model processes this self-provided information to automatically determine data types without requiring external expert intervention, thus resolving the contradiction between maintaining accuracy and improving productivity.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical system of manual expert labeling with an automated machine learning-based classification system. The ML model processes database component information and automatically predicts data types, substituting human expert manual work with an automated computational process that maintains accuracy while significantly improving productivity.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Productivity

If automated methods with limited metadata are used to determine data types, then productivity is improved, but measurement precision deteriorates

Engineering Contradiction:
ImproveproductivityVSAvoidaccuracy of data type classification
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent applies universality by designing a machine learning model that can effectively process multiple types of input information (identifiers, descriptions, metadata) from database components. This multi-functional approach allows the system to achieve accurate data type classification using automated methods with limited metadata, resolving the contradiction between improved productivity and maintained measurement precision.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system employs parameter changes by transforming various database component attributes (identifiers, descriptions, metadata) into features that the machine learning model can process. This transformation enables the model to accurately predict data types even with limited metadata, achieving both high productivity and measurement precision simultaneously.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS11720533B2Automated classification of data types for databases
Publication Date: 2023.08.08 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US11720533B2 patent drawing
  • US11720533B2 patent drawing
  • US11720533B2 patent drawing

AI summary

Techniques for automatically determining different data types found in databases are disclosed. In one example, a computer implemented method comprises receiving a portion of identifying information for one or more components of a database, and generating one or more descriptions for the one or more components based at least in part on the portion of the identifying information for the one or more components. The one or more descriptions are inputted to one or more machine learning models, and, using the one or more machine learning models, one or more data types associated with the one or more components are predicted. The prediction is based at least in part on the one or more descriptions.