Ontology Layer for Querying Heterogeneous Datasets
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Large data repositories comprising multiple datasets in different formats make it difficult for users to query and interact with distributed data effectively, as the datasets are not easily searchable or navigable due to lacking explicit relationships and metadata.
Innovation Solution
A system and method that generates an ontology layer to provide metadata and object graphs, allowing users to define relationships and global properties across datasets, enabling efficient querying and visualization of data through an object view.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If datasets are stored in different formats without explicit relationships, then data diversity and flexibility are improved, but data query efficiency and navigation ease deteriorate
Solution Approach 1:
The patent introduces a metadata layer as an intermediary between the diverse datasets and the user query interface. This metadata layer contains schema information, data type definitions, and relationship descriptions that enable efficient querying without requiring direct access to the heterogeneous data formats. The metadata acts as a mediator that translates user queries into appropriate data access operations across multiple formats.
Solution Approach 2:
The system segments the data management functionality into separate components: the raw datasets remain in their original formats, while a separate metadata layer handles query processing. This segmentation allows each component to maintain its independence and characteristics while working together as a unified system, preserving data diversity without compromising query efficiency.
2Quantity of substance
If datasets are distributed across multiple sources, then data coverage and comprehensiveness are improved, but data integration complexity and processing difficulty increase
Solution Approach 1:
The metadata layer serves multiple functions simultaneously: it describes data schemas, defines relationships between datasets, guides query processing, and enables data integration. This universal component handles all aspects of data management across distributed sources without requiring separate mechanisms for each function, thereby reducing overall system complexity despite increased data coverage.
3Adaptability or versatility
If explicit relationships between datasets are not defined, then data independence and flexibility are improved, but data relationship discovery and analysis difficulty increase
Solution Approach 1:
The system performs preliminary action by pre-defining relationships and schemas in the metadata layer before queries are executed. The metadata contains预先 established relationship information, data type definitions, and structural descriptions that enable the system to understand and navigate data relationships without requiring real-time discovery or complex analysis during query processing.
Data Source
AI summary
This disclosure relates to a system and method for data analysis. According to a first aspect, there is described a method, the method being performed using one or more processors, comprising: receiving one or more user inputs indicative of one or more relationships between data in a plurality of datasets; determining, based on the one or more user inputs, at least one object view for visualizing the data in the plurality of datasets; generating, based on the one or more user inputs, metadata comprising: an object graph indicative of the one or more relationships between two or more of the plurality of datasets; and information identifying the at least one object view; and in response to a query relating to the plurality of datasets, using the metadata to determine how response data responding to the query should be provided.


