Dataset Catalog Search Using Relationship Graphs
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data storage systems face challenges in identifying relevant datasets and understanding their relationships, making it difficult for users to formulate queries effectively, especially as the number of dataset groups grows, with most efforts focused on query formulation rather than dataset identification and relationship analysis.
Innovation Solution
A system that uses human input or machine learning to understand dataset structures and relationships, allowing users to provide examples of data and relationships to search for similar information, thereby identifying relevant datasets and filtering them based on domain overlaps and connection strengths.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If full text search techniques are applied to datasets, then any type of data can be easily handled, but the ability to understand and utilize data structure and relationships is lost
Solution Approach 1:
The patent segments the search process into two distinct components: (1) structure-aware component that analyzes data schemas, relationships, and metadata to understand data organization, and (2) content search component that performs full-text searching. This segmentation allows the system to leverage both full-text versatility and structural understanding without compromise.
Solution Approach 2:
The patent introduces an intermediary layer (data catalog with schema information, relationship graphs, and metadata) that sits between the full-text search engine and the user query. This intermediary enriches full-text search results by adding structural context, data type information, and relationship awareness, thereby preserving structural understanding while maintaining full-text versatility.
2Measurement precision
If data is organized into a well-understand semantic model prior to searching, then powerful search capabilities are achieved, but the difficulty of organizing all data in advance increases
Solution Approach 1:
The patent performs preliminary organization actions automatically during data ingestion and storage, rather than requiring manual pre-organization before searching. The system automatically builds data catalogs, extracts schemas, identifies relationships, and creates metadata structures as data is loaded, enabling powerful semantic search without burdening users with complex pre-organization tasks.
Solution Approach 2:
The patent implements self-service automated processes that independently organize data into semantic models without human intervention. The system automatically analyzes data structures, infers relationships, generates metadata, and builds searchable indexes, thereby achieving high search precision while eliminating the complexity of manual data organization for users.
3Quantity of substance
If the number of available dataset groups grows, then more data is available to answer questions, but the complexity of identifying datasets and their relationships increases
Solution Approach 1:
The patent introduces a data catalog as an intermediary structure that indexes and organizes all available datasets, their schemas, relationships, and metadata. This catalog serves as a navigation layer that enables users to efficiently identify relevant datasets among growing numbers without manually analyzing each dataset's structure, thereby scaling data availability while controlling identification complexity.
Solution Approach 2:
The patent implements feedback mechanisms where the system learns from user search patterns, dataset access frequencies, and query outcomes to automatically refine and optimize the data catalog organization. This feedback-driven optimization enables the system to adapt to growing data volumes by automatically improving dataset identification and relationship discovery without increasing user-facing complexity.
Data Source
Figure 1~3
Figure 2
Figure 4
AI summary
In one embodiment, datasets are stored in a catalog. The datasets are enriched by establishing relationships among the domains in different datasets. A user searches for relevant datasets by providing examples of the domains of interest. The system identifies datasets corresponding to the user-provided examples. The system them identifies connected subsets of the datasets that are directly linked or indirectly linked through other domains. The user provides known relationship examples to filter the connected subsets and to identify the connected subsets that are most relevant to the user's query. The selected connected subsets may be further analyzed by business intelligence/analytics to create pivot tables or to process the data.