Automated Dataset Discovery Engine for Analytics Lifecycle
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional data analytics solutions face challenges in identifying relevant datasets for complex data problems due to increasing data sizes and varieties, making it difficult for business users and data scientists to execute data analytics efficiently.
Innovation Solution
A method and system that automate the data analytics lifecycle by discovering and cataloging datasets through a dataset discovery engine, which includes a data crawling agent and feature extractor to identify and filter relevant datasets for hypothesis testing, while ensuring data privacy and quality.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If conventional data analytics solutions are used, then data analysis can be performed, but the ability to determine relevant datasets among increasing sizes and varieties of data sets is limited
Solution Approach 1:
The patent replaces manual dataset identification and selection processes with an automated discovery system that uses machine learning models, natural language processing, and graph-based algorithms to automatically identify relevant datasets based on hypotheses and data problems
Solution Approach 2:
The patent introduces a dataset discovery service as an intermediary layer between the data analytics system and the vast data lake, which acts as a mediator to automatically identify, filter, and recommend relevant datasets based on analytical requirements without direct human intervention
2Productivity
If manual dataset identification is used, then data scientists can select datasets, but the process is time-consuming and inefficient
Solution Approach 1:
The patent performs preliminary actions by pre-processing and indexing data metadata, creating data catalogs and ontologies in advance, and pre-computing dataset relationships and characteristics so that when a data analysis task arises, the discovery system can quickly retrieve and recommend relevant datasets without time-consuming manual search
Solution Approach 2:
The patent enables the system to serve itself by implementing automated dataset discovery and selection capabilities that do not require human data scientists to manually identify relevant datasets, allowing the system to autonomously perform dataset identification based on analytical requirements and hypotheses
3Ease of operation
If automated dataset discovery is implemented, then dataset identification efficiency improves, but system complexity increases
Solution Approach 1:
The patent segments the complex dataset discovery system into distinct modular components including a data crawling agent for data collection, a feature extractor for metadata processing, a graph generator for relationship modeling, and a machine learning model for dataset recommendation, allowing each component to be independently developed, maintained, and optimized
Data Source
AI summary
An initial work package is obtained. The initial work package defines at least one hypothesis associated with a given data problem, and is generated in accordance with one or more phases of an automated data analytics lifecycle. A plurality of datasets is identified. One or more datasets in the plurality of datasets that are relevant to the at least one hypothesis are discovered. The at least one hypothesis is tested using at least a portion of the one or more discovered datasets.


