Centralized Data Repository for Domain-Specific Database Preparation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The challenge in creating a database for domain-specific applications is the heterogeneity of data sources, which leads to difficulties in associating data across different databases due to varying identifiers for the same entity type, resulting in isolated 'data islands' that restrict the capabilities and business value of applications.
Innovation Solution
A method that involves creating a centralized data repository, identifying pivotal entity types, determining mappings between different identifiers, and selecting subsets of data units to form a reference set and non-reference set, enabling the creation of a database that can seamlessly combine data from diverse sources.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If data from different sources is consolidated into a centralized repository, then data availability and application versatility are improved, but data heterogeneity and identifier variability increase, making data association difficult
Solution Approach 1:
The patent introduces an intermediary layer (data mapping service, identifier resolution layer) that sits between the heterogeneous data sources and the application layer. This intermediary automatically resolves different identifiers for the same entity by maintaining mapping relationships and resolving them to a canonical form, thereby enabling data association without exposing the complexity to applications.
Solution Approach 2:
The patent transforms the identifier parameter from its original heterogeneous forms (different formats, schemas, and semantics across sources) into a unified canonical identifier format. This parameter transformation enables consistent data association across diverse sources while preserving the original identifier information through mapping relationships.
2Reliability
If comprehensive data exploration and consolidation processes are performed manually, then data quality and governance are improved, but development time and resource consumption increase
Solution Approach 1:
The patent performs data exploration, consolidation, and mapping establishment as preliminary actions during the data preparation phase, before applications are developed. By pre-processing the data and establishing identifier mappings in advance, the system eliminates the need for manual data preparation during application development, significantly reducing development time while maintaining data quality.
Solution Approach 2:
The system implements automated self-service capabilities where the data platform automatically performs data exploration, identifies relationships between data sources, establishes mappings, and resolves identifiers without requiring manual intervention. This automation maintains data quality through systematic processing while dramatically reducing the time and resources needed compared to manual approaches.
3Adaptability or versatility
If data is stored in isolated databases with different identifiers, then data source independence is maintained, but application capabilities are restricted due to inability to associate data across sources
Solution Approach 1:
The patent creates a universal identifier resolution mechanism that works across all data sources simultaneously. The mapping service provides multi-functional capabilities: it resolves identifiers, maintains mappings, and enables data association across diverse sources while preserving each source's independence. This universal approach eliminates data silos and enables rich application capabilities without losing data association information.
Data Source
AI summary
A method for creating a database for a domain specific application includes providing a centralized data repository comprising data from different sources, identifying a set of data units of the repository that represent a specific domain, determining a pivotal entity type for an application, determining a mapping between different identifiers of the pivotal entity type, creating a reference set by selecting a first subset of the set of data units using the mapping, wherein the first subset represents the pivotal entity type, selecting, based at least in part on the mapping, a second subset of the set of data units, wherein the second subset represents non-pivotal entity types which are related to instances of the pivotal entity type in the reference set; and creating a database from data units and associated attributes selected from the reference set of data units and the second subset of data units.


