Knowledge Graph for Multi-Source Dataset Discovery
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data curation systems lack the ability to efficiently identify and combine relevant datasets for data-driven insights, as they primarily focus on structural attributes and licensing rights, neglecting the quality and relevance of dataset content, and fail to leverage expert knowledge and semantic search techniques to optimize search results.
Innovation Solution
A system that organizes and displays dataset metadata within a knowledge graph framework, using hybrid quality assurance processes and indexing to enhance metadata attributes, allowing users to find and match multiple linked datasets based on relevance, geographical, temporal, and technical criteria, and provides a comparison matrix for dataset selection.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If traditional relational database structures are used to organize datasets, then data storage and basic retrieval are simplified, but the ability to index and link datasets across multiple sources and subject matter areas is inhibited
Solution Approach 1:
The patent introduces a knowledge graph as an intermediary layer between traditional relational databases and dataset metadata. This knowledge graph uses graph database technology with nodes representing datasets, subjects, and attributes, and edges representing relationships between them. The knowledge graph mediates between the structured storage of relational databases and the need for flexible, multi-dimensional indexing and linking capabilities, enabling users to traverse relationships across datasets from multiple sources without changing the underlying database structure.
2Reliability
If dataset curation focuses on structural attributes and licensing rights, then data organization and legal compliance are improved, but the utility and usability of dataset content to basic crowdsourced keyword tags is minimized
Solution Approach 1:
The patent segments dataset metadata into multiple hierarchical levels: structural attributes (format, size, licensing), content attributes (subject matter, keywords, descriptions), and quality attributes (verification status, expert ratings). This segmentation allows the system to maintain rigorous structural and legal compliance while simultaneously preserving and enhancing content relevance information through separate, dedicated metadata fields that can be independently curated and searched.
Solution Approach 2:
The system performs preliminary quality assurance and expert verification of dataset content before making datasets available for search and selection. Expert curators pre-validate dataset content, add detailed subject matter tags, and assess relevance to various domains in advance. This preliminary action ensures that when users search for datasets, the content relevance information is already prepared and accurate, reducing the need for users to perform additional validation.
3Loss of information
If multiple datasets with interconnected content are needed for data projects, then comprehensive insight and detail are improved, but the complexity of identifying compatible dataset combinations increases
Solution Approach 1:
The system provides feedback mechanisms that analyze user search patterns, dataset usage patterns, and compatibility relationships between datasets. This feedback is used to automatically generate and update recommendations for dataset combinations, learn from successful multi-dataset projects, and suggest compatible datasets that users may not have considered. The feedback loop continuously improves the system's ability to identify optimal dataset combinations for comprehensive data insights.
Solution Approach 2:
The knowledge graph serves as an intermediary that automatically identifies and presents compatible dataset combinations to users. Instead of requiring users to manually search for and evaluate multiple datasets for compatibility, the knowledge graph's graph database technology automatically traverses relationships between datasets, subjects, and attributes to identify coherent combinations that meet project requirements. This intermediary processing significantly reduces the complexity of identifying compatible multi-dataset configurations.
4Reliability
If extensive data verification and quality assurance processes are implemented, then data quality and authenticity are improved, but the time and resources required for dataset curation increase
Solution Approach 1:
The system implements partial verification processes where not all datasets require the same level of quality assurance. Datasets are categorized by risk level, source reliability, and intended use, with corresponding verification requirements. High-risk datasets receive extensive expert review and validation, while low-risk datasets from trusted sources undergo streamlined verification. This partial action approach maintains data authenticity where needed while reducing curation time for lower-risk datasets.
Data Source
AI summary
A graph-based data cataloging system, product and method that structures expert knowledge and statistically driven data analytics into a system-based framework for finding and relating enhanced metadata on subject-relevant, curated datasets from disparate, externally held data sources is shown. Displayed across a knowledge graph of nodes of datasets linked by their metadata attributes, the system simplifies the search and retrieval of multiple datasets of relevance to a user's technical, content, and resource-driven needs.


