Reconciled Data Storage System for Heterogeneous Dataset Integration
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The integration of input datasets from heterogeneous sources with different structures and formats is challenging due to the need for manual interaction and interpretation, and existing technologies impose processing overheads in data centers, particularly in cross-device traversals and access operations.
Innovation Solution
A data storage system with a processor and memory that includes a reconciled data store and a registry, where each registered entity is stored with property-value pairs from multiple datasets, and identifier property-value pairs are used to consolidate data from different sources, reducing manual intervention and enhancing machine recognition of context and knowledge domains.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual interaction and interpretation are used to integrate datasets from heterogeneous sources, then data integration accuracy is improved, but processing time and labor cost increase
Solution Approach 1:
The system enables automated self-service through machine learning models that automatically recognize entities, interpret data context, and integrate datasets without manual intervention. The models learn from training data to autonomously perform data reconciliation tasks, eliminating the need for human operators while maintaining high accuracy in identifying and merging data from heterogeneous sources.
Solution Approach 2:
Manual mechanical processes of data integration are replaced with automated computational systems. Machine learning algorithms substitute human analysts by automatically performing entity recognition, data interpretation, and integration tasks that previously required manual expertise and time-consuming operations.
2Measurement precision
If manual interpretation is used to understand input datasets, then data interpretation accuracy is improved, but productivity decreases
Solution Approach 1:
The system performs automated self-service through trained machine learning models that independently interpret datasets by recognizing entities, understanding data context, and determining relationships between data elements. This automation maintains high interpretation accuracy while enabling parallel processing of multiple datasets simultaneously, dramatically increasing productivity.
Solution Approach 2:
Machine learning models are pre-trained on extensive datasets before deployment, performing preliminary learning of data patterns, entities, and relationships. This preliminary action enables the models to quickly and accurately interpret new datasets without requiring manual expertise during actual processing, thereby maintaining accuracy while scaling productivity.
3Ease of operation
If cross-device traversals and access operations are performed in data centres, then data access capability is improved, but processing overhead increases
Solution Approach 1:
The system merges data from multiple heterogeneous sources into a unified reconciled dataset, eliminating the need for repeated cross-device traversals. By consolidating data access operations and maintaining a centralized reconciled data store, the system reduces processing overhead and energy consumption while preserving comprehensive data access capability.
Solution Approach 2:
Data reconciliation and integration are performed in advance, creating a pre-processed unified dataset that can be accessed without repeated traversal operations. This preliminary action of consolidating data eliminates subsequent overhead from multiple access operations across different data sources.
Data Source
Figure 1~2
Figure 3
Figure 4
AI summary
Embodiments include: a data storage system comprising a processor coupled to a memory, the memory including a reconciled data store and a registry, the processor being configured to execute a process comprising: - in the reconciled data store, storing, with respect to each of a plurality of registered entities: an array of registered entity property-value pairs representing the registered entity, from each dataset from among a plurality of (heterogeneous) datasets, wherein each of the registered entity data property-value pairs comprises a property label representing a property of a registered entity and a value representing a value range of the property for the respective registered entity; - in a registry, storing, with respect to each of the plurality of registered entities: a registry entry comprising at least one identifier property-value pair with respect to each of the plurality of datasets, wherein each of the identifier property-value pairs comprises an identifier property label representing an identifier property of a registered entity, uniquely identifying the registered entity (within the respective dataset), and an identifier value representing a value of the identifier property for the respective registered entity; - acquiring a dataset for reconciliation with the reconciled data store, the acquired dataset including a plurality of acquired dataset property-value pairs for each of a first set of one or more acquired dataset entities, wherein each of the acquired dataset data property-value pairs comprises a property label representing a property of an acquired dataset entity and a value representing a value range of the property for the respective acquired dataset entity. The process further comprises, for each of the one or more acquired dataset entities: - identifying an identifier property-value pair stored in the registry matching an acquired dataset property-value pair for the acquired dataset entity; and - consolidating the acquired dataset property-value pairs for the acquired dataset entity into the array of registered entity property-value pairs stored with respect to the registered entity identified by the identifier value of the identified identifier property-value pair.