Dataset Catalog Search Using Relationship Graphs

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing data storage systems face challenges in identifying relevant datasets and understanding their relationships, making it difficult for users to formulate queries effectively, especially as the number of dataset groups grows, with most efforts focused on query formulation rather than dataset identification and relationship analysis.

Innovation Solution

A system that uses human input or machine learning to understand dataset structures and relationships, allowing users to provide examples of data and relationships to search for similar information, thereby identifying relevant datasets and filtering them based on domain overlaps and connection strengths.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If full text search techniques are applied to datasets, then any type of data can be easily handled, but the ability to understand and utilize data structure and relationships is lost

Engineering Contradiction:
Improveability to handle any type of dataVSAvoidinability to understand data structure and relationships
Core Design Contradiction:
Adaptability or versatilityVSLoss of information

Solution Approach 1:

The patent segments the search process into two distinct components: (1) structure-aware component that analyzes data schemas, relationships, and metadata to understand data organization, and (2) content search component that performs full-text searching. This segmentation allows the system to leverage both full-text versatility and structural understanding without compromise.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary layer (data catalog with schema information, relationship graphs, and metadata) that sits between the full-text search engine and the user query. This intermediary enriches full-text search results by adding structural context, data type information, and relationship awareness, thereby preserving structural understanding while maintaining full-text versatility.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If data is organized into a well-understand semantic model prior to searching, then powerful search capabilities are achieved, but the difficulty of organizing all data in advance increases

Engineering Contradiction:
Improvesearch capability precisionVSAvoiddifficulty of organizing data in advance
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent performs preliminary organization actions automatically during data ingestion and storage, rather than requiring manual pre-organization before searching. The system automatically builds data catalogs, extracts schemas, identifies relationships, and creates metadata structures as data is loaded, enabling powerful semantic search without burdening users with complex pre-organization tasks.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements self-service automated processes that independently organize data into semantic models without human intervention. The system automatically analyzes data structures, infers relationships, generates metadata, and builds searchable indexes, thereby achieving high search precision while eliminating the complexity of manual data organization for users.

Inventive Principle:
Principle #25Self-service

3Quantity of substance

If the number of available dataset groups grows, then more data is available to answer questions, but the complexity of identifying datasets and their relationships increases

Engineering Contradiction:
Improveamount of available dataVSAvoidcomplexity of identifying datasets and relationships
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The patent introduces a data catalog as an intermediary structure that indexes and organizes all available datasets, their schemas, relationships, and metadata. This catalog serves as a navigation layer that enables users to efficiently identify relevant datasets among growing numbers without manually analyzing each dataset's structure, thereby scaling data availability while controlling identification complexity.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent implements feedback mechanisms where the system learns from user search patterns, dataset access frequencies, and query outcomes to automatically refine and optimize the data catalog organization. This feedback-driven optimization enables the system to adapt to growing data volumes by automatically improving dataset identification and relationship discovery without increasing user-facing complexity.

Inventive Principle:
Principle #23Feedback

Data Source

PatentEP2836935B1Finding data in connected corpuses using examples
Publication Date: 2019.06.05 MICROSOFT TECHNOLOGY LICENSING LLC
  • EP2836935B1 patent drawingFigure 1~3
  • EP2836935B1 patent drawingFigure 2
  • EP2836935B1 patent drawingFigure 4

AI summary

In one embodiment, datasets are stored in a catalog. The datasets are enriched by establishing relationships among the domains in different datasets. A user searches for relevant datasets by providing examples of the domains of interest. The system identifies datasets corresponding to the user-provided examples. The system them identifies connected subsets of the datasets that are directly linked or indirectly linked through other domains. The user provides known relationship examples to filter the connected subsets and to identify the connected subsets that are most relevant to the user's query. The selected connected subsets may be further analyzed by business intelligence/analytics to create pivot tables or to process the data.