RDF Data Storage in Relational Hash Tables
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional RDF database query processing is inefficient due to the need for numerous joins and self-joins, which are slow and difficult to optimize, especially when dealing with schema-less data in relational databases, where predicates for subjects vary widely and storage in a single property table is not feasible.
Innovation Solution
A hash table is created with each subject represented in a row, using a hash function to insert predicate/object data in pairs of columns, allowing for efficient storage and query processing by minimizing self-joins and optimizing space usage, while maintaining compatibility with existing relational database systems.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If RDF data is stored in a conventional triple store format, then the data can be stored with a natural graph structure, but query processing becomes slow due to numerous joins and self-joins
Solution Approach 1:
The patent segments the RDF triple store into multiple relational tables with specific purposes: a subjects table, a predicates table, and a join table. This segmentation allows queries to be processed through optimized relational joins rather than complex self-joins on a single triple table, significantly improving query performance while maintaining the ability to store and represent the complete graph structure.
Solution Approach 2:
The patent introduces an intermediary join table that connects the subjects table and predicates table. This intermediary structure enables efficient relational joins by providing a standardized interface between subject and predicate data, eliminating the need for complex self-joins on the original triple store while preserving the natural graph structure representation.
2Quantity of substance
If a single property table is used to store all predicates for each subject, then storage space is optimized, but the table becomes too complex and difficult to manage when predicates vary widely
Solution Approach 1:
The patent segments the single complex property table into multiple specialized tables: a subjects table storing subject identifiers and basic information, a predicates table storing predicate definitions and metadata, and a join table linking subjects to their predicates. This segmentation reduces the complexity of any single table while maintaining storage efficiency through normalized relational structure.
Solution Approach 2:
The patent transforms the two-dimensional single table structure into a three-dimensional normalized schema with separate tables for subjects, predicates, and their relationships. This dimensional change allows the system to handle widely varying predicates efficiently by storing them in a dedicated predicates table rather than creating sparse columns in a single table.
3Adaptability or versatility
If multiple joins and self-joins are used to retrieve related records, then complete graph relationships can be accessed, but query processing time increases significantly
Solution Approach 1:
The patent segments the query processing workload by organizing data into specialized tables that can be joined efficiently. The subjects table, predicates table, and join table are designed to minimize the number and complexity of joins required, allowing graph relationships to be retrieved through optimized relational operations rather than complex self-joins.
Solution Approach 2:
The patent performs preliminary organization of RDF data into normalized relational tables during the data loading phase. By pre-establishing the relationships between subjects and predicates in a structured format with proper foreign keys and indexes, the system eliminates the need for complex runtime joins, significantly reducing query processing time while maintaining full graph relationship retrieval capability.
Data Source
AI summary
A method (and structure) of storing schema-less data of a dataset in a relational database, includes constructing a hash table for the schema-less data, using a processor on a computer. Data in the dataset is stored in a tuple format including a subject along with at least one other entity associated to the subject. Each row of the hashtable will be dedicated to a subject of the dataset, and at least one of the at least one other entity associated with the subject in the row is to be stored in a pair-wise manner in that row of the hashtable. In an exemplary embodiment, RDF data that uses triples (subject, predicate, object) is stored with the predicate/object stored in the pair-wise manner in its associated subject row.


