Automated Dataset Discovery Engine for Analytics Lifecycle

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional data analytics solutions face challenges in identifying relevant datasets for complex data problems due to increasing data sizes and varieties, making it difficult for business users and data scientists to execute data analytics efficiently.

Innovation Solution

A method and system that automate the data analytics lifecycle by discovering and cataloging datasets through a dataset discovery engine, which includes a data crawling agent and feature extractor to identify and filter relevant datasets for hypothesis testing, while ensuring data privacy and quality.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If conventional data analytics solutions are used, then data analysis can be performed, but the ability to determine relevant datasets among increasing sizes and varieties of data sets is limited

Engineering Contradiction:
Improveability to determine relevant datasetsVSAvoidsizes and varieties of data sets
Core Design Contradiction:
Adaptability or versatilityVSQuantity of substance

Solution Approach 1:

The patent replaces manual dataset identification and selection processes with an automated discovery system that uses machine learning models, natural language processing, and graph-based algorithms to automatically identify relevant datasets based on hypotheses and data problems

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent introduces a dataset discovery service as an intermediary layer between the data analytics system and the vast data lake, which acts as a mediator to automatically identify, filter, and recommend relevant datasets based on analytical requirements without direct human intervention

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If manual dataset identification is used, then data scientists can select datasets, but the process is time-consuming and inefficient

Engineering Contradiction:
Improveefficiency of data analytics executionVSAvoidtime to identify relevant datasets
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent performs preliminary actions by pre-processing and indexing data metadata, creating data catalogs and ontologies in advance, and pre-computing dataset relationships and characteristics so that when a data analysis task arises, the discovery system can quickly retrieve and recommend relevant datasets without time-consuming manual search

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent enables the system to serve itself by implementing automated dataset discovery and selection capabilities that do not require human data scientists to manually identify relevant datasets, allowing the system to autonomously perform dataset identification based on analytical requirements and hypotheses

Inventive Principle:
Principle #25Self-service

3Ease of operation

If automated dataset discovery is implemented, then dataset identification efficiency improves, but system complexity increases

Engineering Contradiction:
Improveease of executing data analyticsVSAvoidcomplexity of dataset discovery system
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent segments the complex dataset discovery system into distinct modular components including a data crawling agent for data collection, a feature extractor for metadata processing, a graph generator for relationship modeling, and a machine learning model for dataset recommendation, allowing each component to be independently developed, maintained, and optimized

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS9235630B1Dataset discovery in data analytics
Publication Date: 2016.01.12 EMC IP HLDG CO LLC
  • US9235630B1 patent drawing
  • US9235630B1 patent drawing
  • US9235630B1 patent drawing

AI summary

An initial work package is obtained. The initial work package defines at least one hypothesis associated with a given data problem, and is generated in accordance with one or more phases of an automated data analytics lifecycle. A plurality of datasets is identified. One or more datasets in the plurality of datasets that are relevant to the at least one hypothesis are discovered. The at least one hypothesis is tested using at least a portion of the one or more discovered datasets.