Intelligent Data Management System for Context-Aware Dataset Discovery

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for business users to access and discover data are inefficient, requiring knowledge of data categorization and labels, leading to less relevant data and missed exploration opportunities, as they typically rely on manual or automatic labeling without context, making it difficult to navigate and utilize relevant datasets effectively.

Innovation Solution

The implementation of an Intelligent Data Management System (IDMS) that uses question-based dataset generation, employing metadata such as business questions, intent, and historical data to identify and suggest relevant datasets, allowing users to input their problems rather than specific data queries, and utilizing AI-powered knowledge bases to guide data exploration and selection.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If users search for data using known labels or tags, then data access is enabled, but the user must be fully aware of the probable categorization of data which is inefficient and results in loss of data exploration opportunities

Engineering Contradiction:
Improvedata accessVSAvoiddata exploration time
Core Design Contradiction:
Ease of operationVSLoss of time

Solution Approach 1:

The patent introduces an intermediary system (data exploration system with AI/ML models) that mediates between the user's natural language query and the labeled data repository. This intermediary automatically interprets user intent, maps it to relevant data categories, and retrieves appropriate datasets without requiring users to know the labeling schema, thus resolving the contradiction between ease of access and exploration efficiency

Inventive Principle:
Principle #24Intermediary (Mediator)

2Ease of operation

If users search for data using known labels or tags, then data access is enabled, but this approach produces less relevant data which may adversely affect decisions

Engineering Contradiction:
Improvedata query simplicityVSAvoiddata relevance
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The patent changes the query parameter from structured labels/tags to natural language questions. The system transforms user questions into meaningful data queries using AI/ML techniques, allowing users to express their information needs in everyday language while the system automatically identifies relevant labeled data, thereby improving both query simplicity and data relevance simultaneously

Inventive Principle:
Principle #35Parameter changes

3Productivity

If manual or automatic labeling is used without context, then data categorization is achieved, but it becomes difficult to navigate and utilize relevant datasets effectively

Engineering Contradiction:
Improvedata categorization efficiencyVSAvoiddata navigation
Core Design Contradiction:
ProductivityVSEase of operation

Solution Approach 1:

The patent applies preliminary action by pre-computing and storing contextual relationships between data labels, tags, and semantic meanings in the background. When users navigate or query data, the system leverages these pre-computed contextual mappings to quickly present relevant datasets without requiring users to manually navigate through categorized data structures, thus improving both categorization efficiency and navigation ease

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11860943B2Method of “outcome driven data exploration” for datasets, business questions, and pipelines based on similarity mapping of business needs and asset use overlap
Publication Date: 2024.01.02 EMC IP HLDG CO LLC
  • US11860943B2 patent drawing
  • US11860943B2 patent drawing
  • US11860943B2 patent drawing

AI summary

One example method includes receiving a query that recites a particular question for which a user who originated the query needs an answer, parsing the query to identify the question, identifying information that is responsive to the question, presenting the information to the user in a user-selectable form, and receiving, from the user, a selection of the information. In some cases, the information presented to the user may include one or more datasets, or one or more pipelines.