Ontology-Based Query Suggestion for Data Repositories

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Large-scale data analytic systems face challenges in efficiently querying and interacting with multiple disparate datasets, particularly for non-technical users who need to construct sophisticated queries but lack the necessary knowledge.

Innovation Solution

A method and system that utilize an ontology to suggest database queries in natural language, allowing users to input keywords, determine relevant datasets, identify metadata relationships, and construct object views based on selected queries.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If relational data is stored within the datasets themselves, then the datasets can be easily queried and joined, but the size of the datasets increases significantly

Engineering Contradiction:
Improvequery capabilityVSAvoiddataset size
Core Design Contradiction:
Ease of operationVSQuantity of substance

Solution Approach 1:

The patent introduces an ontology layer as an intermediary between the data repository and users. The ontology contains metadata that describes relationships between datasets without duplicating the actual data. This mediator enables querying and joining capabilities while keeping the original datasets compact, as the relationship information is stored separately in the ontology rather than within each dataset.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If data is stored in a structured relational way, then analytical processing can be performed efficiently, but non-technical users find it difficult to interact with and construct queries

Engineering Contradiction:
Improveanalytic processing efficiencyVSAvoiduser interaction difficulty
Core Design Contradiction:
ProductivityVSEase of operation

Solution Approach 1:

The ontology serves as an intermediary that translates between the structured relational data format and natural language queries. It provides a layer of abstraction that allows non-technical users to interact with the data using everyday language while the system handles the complex relational queries underneath, maintaining analytic processing efficiency without increasing user complexity.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The ontology layer provides multiple functions: it stores metadata about datasets, defines relationships between them, enables natural language query translation, and supports both technical and non-technical users. This multi-functional approach allows the system to maintain efficient analytic processing while being accessible to users with varying levels of technical expertise.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Ease of operation

If an ontology layer with metadata is introduced to describe relationships between datasets, then non-technical users can easily query data, but the system complexity increases

Engineering Contradiction:
Improvequery construction easeVSAvoidsystem architecture complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent segments the system into distinct layers: a data repository layer for storing actual datasets, an ontology layer for storing metadata and relationships, and a query processing layer for handling user requests. This segmentation allows each layer to be developed and maintained independently, reducing overall system complexity despite the addition of the ontology layer.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS20250117431A1System and method for querying a data repository
Publication Date: 2025.04.10 PALANTIR TECHNOLOGIES INC
  • US20250117431A1 patent drawing
  • US20250117431A1 patent drawing
  • US20250117431A1 patent drawing

AI summary

The present disclosure relates to methods and systems for querying data in a data repository. According to a first aspect, this disclosure describes a method of querying a database, comprising: receiving, at a computing device, a plurality of keywords; determining, by the computer device, a plurality of datasets relating to the keywords; identifying, by the computer device, metadata for the plurality of datasets indicating a relationship between the datasets by examining an ontology associated with the datasets; providing, by the computer device, one or more suggested database queries in natural language form, the one or more suggested database queries constructed based on the plurality of keywords and the metadata; receiving, by the computing device, a selection of the one or more suggested database queries; and constructing, by the computer device, an object view for the plurality of datasets based on the selected query and the metadata.