Linked Data Model for Heterogeneous Database Integration
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Traditional data management approaches are inflexible and struggle to efficiently link and manage diverse data systems, particularly in complex environments like IoT settings, where content heterogeneity and scalability challenges hinder effective data control and integration across disparate databases.
Innovation Solution
A data control system that utilizes a linked data model (LDM) with a domain knowledge graph and system metadata to abstract data sources and applications, enabling efficient data ingestion, consumption, and exploration by maintaining relational linkages between data objects across multiple databases, and facilitating onboarding of new data sources through core models.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If traditional batch-driven ETL processes are used for data management, then data integration can be achieved, but the system becomes inflexible and struggles with content heterogeneity and scalability
Solution Approach 1:
The patent introduces an intermediary layer (data lake with metadata catalog and query engine) between diverse data sources and applications. This intermediary abstracts the complexity of heterogeneous data systems, enabling flexible data access without requiring complex ETL processes for each data source.
Solution Approach 2:
The data lake architecture provides a universal platform that can handle multiple types of data sources (relational databases, NoSQL databases, file systems, IoT devices) through a common interface, eliminating the need for source-specific integration logic and improving system flexibility.
2Productivity
If data is stored across multiple disparate databases, then data diversity is maintained, but efficient linking and retrieval of data objects becomes difficult
Solution Approach 1:
The patent uses a metadata catalog as an intermediary that stores relational linkage information between data objects across different databases. This metadata layer enables efficient querying and retrieval of related data without requiring direct access to multiple disparate database systems.
Solution Approach 2:
The system creates copies of relational linkage information in the metadata catalog, allowing applications to query data relationships without accessing the actual source databases. This copying mechanism preserves linkage information while enabling efficient data retrieval.
3Reliability
If traditional data control approaches are used in IoT environments, then basic data storage is achieved, but robustness and scalability of ETL processes are compromised
Solution Approach 1:
The patent segments the data control architecture into independent components: data ingestion layer, data lake storage, metadata catalog, and query engine. This segmentation allows each component to scale independently and improves robustness by isolating failures to specific segments without affecting the entire system.
Solution Approach 2:
The system employs dynamic data ingestion capabilities that can adapt to new data sources and formats in real-time, allowing the ETL process to scale and evolve with IoT environments without requiring complete system redesign.
Data Source
AI summary
A system creates an abstraction layer surrounding a diverse data system including multiple different databases. Data is received from data sources and ingested into the various databases according to a core model. New instances of the core model are created and added to a larger linked data model (LDM) when new data sources are added to the system. The LDM captures the linkages between different linked data objects and links across different databases. Accordingly, applications are able to access or explore the linked data stored in different databases without prior knowledge of the linking relationships.


