Intelligent Data Management System for Context-Aware Dataset Discovery
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for business users to access and discover data are inefficient, requiring knowledge of data categorization and labels, leading to less relevant data and missed exploration opportunities, as they typically rely on manual or automatic labeling without context, making it difficult to navigate and utilize relevant datasets effectively.
Innovation Solution
The implementation of an Intelligent Data Management System (IDMS) that uses question-based dataset generation, employing metadata such as business questions, intent, and historical data to identify and suggest relevant datasets, allowing users to input their problems rather than specific data queries, and utilizing AI-powered knowledge bases to guide data exploration and selection.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If users search for data using known labels or tags, then data access is enabled, but the user must be fully aware of the probable categorization of data which is inefficient and results in loss of data exploration opportunities
Solution Approach 1:
The patent introduces an intermediary system (data exploration system with AI/ML models) that mediates between the user's natural language query and the labeled data repository. This intermediary automatically interprets user intent, maps it to relevant data categories, and retrieves appropriate datasets without requiring users to know the labeling schema, thus resolving the contradiction between ease of access and exploration efficiency
2Ease of operation
If users search for data using known labels or tags, then data access is enabled, but this approach produces less relevant data which may adversely affect decisions
Solution Approach 1:
The patent changes the query parameter from structured labels/tags to natural language questions. The system transforms user questions into meaningful data queries using AI/ML techniques, allowing users to express their information needs in everyday language while the system automatically identifies relevant labeled data, thereby improving both query simplicity and data relevance simultaneously
3Productivity
If manual or automatic labeling is used without context, then data categorization is achieved, but it becomes difficult to navigate and utilize relevant datasets effectively
Solution Approach 1:
The patent applies preliminary action by pre-computing and storing contextual relationships between data labels, tags, and semantic meanings in the background. When users navigate or query data, the system leverages these pre-computed contextual mappings to quickly present relevant datasets without requiring users to manually navigate through categorized data structures, thus improving both categorization efficiency and navigation ease
Data Source
AI summary
One example method includes receiving a query that recites a particular question for which a user who originated the query needs an answer, parsing the query to identify the question, identifying information that is responsive to the question, presenting the information to the user in a user-selectable form, and receiving, from the user, a selection of the information. In some cases, the information presented to the user may include one or more datasets, or one or more pipelines.


