Automated Database Data Type Classification via Machine Learning
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for data modernization face challenges in determining data types in legacy databases due to limited metadata and the inefficiency of manual labeling by subject matter experts, which complicates the process of moving data to modern databases.
Innovation Solution
A computer-implemented method using machine learning models to predict data types by generating descriptions from partial database component information, expanding acronyms and abbreviations, and utilizing trained models with labeled data to classify data types automatically.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual labeling by subject matter experts is used to determine data types, then accuracy of data type classification can be maintained, but productivity is reduced and loss of time increases
Solution Approach 1:
The system enables self-service by allowing the database components themselves to provide sufficient information for data type classification through their identifiers, descriptions, and metadata. The machine learning model processes this self-provided information to automatically determine data types without requiring external expert intervention, thus resolving the contradiction between maintaining accuracy and improving productivity.
Solution Approach 2:
The patent replaces the mechanical system of manual expert labeling with an automated machine learning-based classification system. The ML model processes database component information and automatically predicts data types, substituting human expert manual work with an automated computational process that maintains accuracy while significantly improving productivity.
2Measurement precision
If manual labeling by subject matter experts is used to determine data types, then accuracy of data type classification can be maintained, but loss of time increases
Solution Approach 1:
The system enables self-service by allowing the database components themselves to provide sufficient information for data type classification through their identifiers, descriptions, and metadata. The machine learning model processes this self-provided information to automatically determine data types without requiring external expert intervention, thus resolving the contradiction between maintaining accuracy and improving productivity.
Solution Approach 2:
The patent replaces the mechanical system of manual expert labeling with an automated machine learning-based classification system. The ML model processes database component information and automatically predicts data types, substituting human expert manual work with an automated computational process that maintains accuracy while significantly improving productivity.
3Productivity
If automated methods with limited metadata are used to determine data types, then productivity is improved, but measurement precision deteriorates
Solution Approach 1:
The patent applies universality by designing a machine learning model that can effectively process multiple types of input information (identifiers, descriptions, metadata) from database components. This multi-functional approach allows the system to achieve accurate data type classification using automated methods with limited metadata, resolving the contradiction between improved productivity and maintained measurement precision.
Solution Approach 2:
The system employs parameter changes by transforming various database component attributes (identifiers, descriptions, metadata) into features that the machine learning model can process. This transformation enables the model to accurately predict data types even with limited metadata, achieving both high productivity and measurement precision simultaneously.
Data Source
AI summary
Techniques for automatically determining different data types found in databases are disclosed. In one example, a computer implemented method comprises receiving a portion of identifying information for one or more components of a database, and generating one or more descriptions for the one or more components based at least in part on the portion of the identifying information for the one or more components. The one or more descriptions are inputted to one or more machine learning models, and, using the one or more machine learning models, one or more data types associated with the one or more components are predicted. The prediction is based at least in part on the one or more descriptions.


