Linked Data Indexing for Real-Time Data Warehouse Updates
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The existing ETL process for building a data warehouse is time-consuming, complex, and difficult to maintain, especially when new data types are introduced or changes are made to existing data types, leading to stale data and the inability to enforce data control rules automatically.
Innovation Solution
A data indexing system using linked data standards, where servers maintain resource sets with tracked resource sets that describe resources using RDF and URIs, allowing for near real-time data extraction and querying across multiple data sources without the need for transformation, and enabling low-latency querying and incremental updates.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If ETL process is used to build data warehouse, then data can be stored and queried from multiple sources, but the process is time-consuming and data becomes stale
Solution Approach 1:
The patent extracts only the essential data elements and relationships from multiple sources using linked data standards, storing them in a simplified graph structure that can be updated incrementally rather than requiring complete ETL cycles, thus maintaining data freshness while reducing extraction time
Solution Approach 2:
The system performs preliminary data validation and transformation at the source using standardized schemas, so that when data is extracted and loaded, it requires minimal processing. This preliminary structuring allows for faster, near-real-time data warehouse updates without compromising data quality
2Reliability
If ETL process is used to build data warehouse, then data can be centralized, but the process is complex and difficult to maintain
Solution Approach 1:
The patent changes the fundamental parameters of data representation by adopting linked data standards (RDF triples, URIs, schemas) instead of traditional relational formats. This parameter change simplifies the data model and makes the system more adaptable to new data types, reducing maintenance complexity while maintaining centralized data storage
Solution Approach 2:
The system implements a universal data model using linked data standards that can accommodate multiple data types and sources through a common schema framework. This universality allows the same infrastructure to handle diverse data sources without requiring complex, custom ETL processes for each data type, thereby reducing overall system complexity
3Reliability
If traditional data warehouse is built, then data can be stored, but automatic enforcement of data control rules is not possible
Solution Approach 1:
The patent implements feedback mechanisms where data control rules are automatically validated as data is ingested and stored. The system continuously checks incoming data against defined schemas and constraints, providing immediate feedback and rejection of non-compliant data, thereby automatically enforcing data control rules without manual intervention
Data Source
AI summary
A data indexing system including a plurality of servers and a tracked resource set client is provided. Each of the servers includes a plurality of resources that are part of a resource set. Each of the servers also includes a tracked resource set corresponding to the resource set. The tracked resource set describes the plurality of resources located in the resource set. The server identifies the plurality of resources using rules of linked data. The tracked resource set client is in communication with the plurality of servers. The tracked resource set client has a data index. The data index is built and kept up to date using the tracked resource set of each of the plurality of servers.


