Automated Dataset Discovery for Machine Learning Models
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current cognitive systems and automated machine learning platforms require skilled individuals to manually search and integrate external datasets with initial datasets to enhance the quality of machine learning models, which is time-consuming and labor-intensive.
Innovation Solution
A learning-based approach is implemented to automatically discover and join related datasets with initial datasets, using a convolutional neural network to analyze and vectorize data structures, content, and metadata, facilitating the selection of complementary data for improved model generation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If manual data gathering and dataset integration is performed by skilled individuals, then the quality and completeness of the initial dataset can be improved, but the time consumption and labor intensity increase significantly
Solution Approach 1:
The system enables automated self-service through the dataset discovery service that automatically searches, evaluates, and integrates external datasets without requiring manual intervention. The cognitive system autonomously performs data gathering tasks by accessing external data sources, evaluating dataset quality using multiple criteria, and integrating selected datasets into the training process, thereby eliminating the need for skilled individuals to manually perform these time-consuming tasks while maintaining data quality standards
Solution Approach 2:
The patent replaces the mechanical manual process of data gathering and integration with an automated cognitive system. The manual research and exploration of external datasets by skilled individuals is substituted by an automated dataset discovery service that uses machine learning models to search, evaluate, and integrate datasets. This mechanical substitution transforms the labor-intensive manual process into an automated computational process, significantly reducing time consumption while maintaining or improving data quality
2Reliability
If manual research and exploration of external datasets is performed, then the completeness and robustness of the machine learning model can be improved, but the labor intensity and expertise requirements increase
Solution Approach 1:
The cognitive system performs self-service by autonomously conducting the research and exploration of external datasets. The dataset discovery service automatically identifies relevant external data sources, evaluates their quality and relevance using predefined criteria, and integrates selected datasets into the training process without requiring skilled individuals to manually research and explore external datasets. This maintains model robustness while eliminating the need for specialized expertise in manual data exploration
Solution Approach 2:
The patent introduces a dataset discovery service as an intermediary between the machine learning training process and external data sources. This intermediary automatically handles the complex tasks of searching, evaluating, and integrating external datasets, shielding users from the complexity of manual data exploration. The intermediary service translates high-level training requirements into automated data gathering and integration operations, maintaining model robustness while simplifying the operational process for users
3Productivity
If automated dataset discovery is implemented using cognitive systems, then productivity and efficiency are improved, but the system complexity increases
Solution Approach 1:
The patent segments the automated dataset discovery system into distinct functional modules: a dataset discovery service that searches and identifies external datasets, an evaluation component that assesses dataset quality using multiple criteria, and an integration component that merges selected datasets with the initial training dataset. This segmentation organizes the complex automated process into manageable, modular components that can be independently developed, maintained, and optimized, thereby improving productivity while managing system complexity through structured modularity
Data Source
AI summary
Embodiments relate to a system, program product, and method for leveraging cognitive systems to facilitate the automated data table discovery for automated machine learning, and, more specifically, to leveraging a trained cognitive system to automatically search for additional data in an external data source that may be merged with an initial user-selected data table to generate a more robust machine learning model. Manual efforts to find and validate data appropriate for building and training a particular model for a particular task are significantly reduced. Specifically, a learning-based approach to leverage with machine learning models to automatically discover related datasets and join the datasets for a given initial dataset is disclosed herein. Operations that include dataset selection facilitate continued reinforcement learning of the systems.


