Ontology-Based Query Suggestion for Data Repositories
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Large-scale data analytic systems face challenges in efficiently querying and interacting with multiple disparate datasets, particularly for non-technical users who need to construct sophisticated queries but lack the necessary knowledge.
Innovation Solution
A method and system that utilize an ontology to suggest database queries in natural language, allowing users to input keywords, determine relevant datasets, identify metadata relationships, and construct object views based on selected queries.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If relational data is stored within the datasets themselves, then the datasets can be easily queried and joined, but the size of the datasets increases significantly
Solution Approach 1:
The patent introduces an ontology layer as an intermediary between the data repository and users. The ontology contains metadata that describes relationships between datasets without duplicating the actual data. This mediator enables querying and joining capabilities while keeping the original datasets compact, as the relationship information is stored separately in the ontology rather than within each dataset.
2Productivity
If data is stored in a structured relational way, then analytical processing can be performed efficiently, but non-technical users find it difficult to interact with and construct queries
Solution Approach 1:
The ontology serves as an intermediary that translates between the structured relational data format and natural language queries. It provides a layer of abstraction that allows non-technical users to interact with the data using everyday language while the system handles the complex relational queries underneath, maintaining analytic processing efficiency without increasing user complexity.
Solution Approach 2:
The ontology layer provides multiple functions: it stores metadata about datasets, defines relationships between them, enables natural language query translation, and supports both technical and non-technical users. This multi-functional approach allows the system to maintain efficient analytic processing while being accessible to users with varying levels of technical expertise.
3Ease of operation
If an ontology layer with metadata is introduced to describe relationships between datasets, then non-technical users can easily query data, but the system complexity increases
Solution Approach 1:
The patent segments the system into distinct layers: a data repository layer for storing actual datasets, an ontology layer for storing metadata and relationships, and a query processing layer for handling user requests. This segmentation allows each layer to be developed and maintained independently, reducing overall system complexity despite the addition of the ontology layer.
Data Source
AI summary
The present disclosure relates to methods and systems for querying data in a data repository. According to a first aspect, this disclosure describes a method of querying a database, comprising: receiving, at a computing device, a plurality of keywords; determining, by the computer device, a plurality of datasets relating to the keywords; identifying, by the computer device, metadata for the plurality of datasets indicating a relationship between the datasets by examining an ontology associated with the datasets; providing, by the computer device, one or more suggested database queries in natural language form, the one or more suggested database queries constructed based on the plurality of keywords and the metadata; receiving, by the computing device, a selection of the one or more suggested database queries; and constructing, by the computer device, an object view for the plurality of datasets based on the selected query and the metadata.


