Unified Data Integration System for Heterogeneous Database Schemas

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing systems face challenges in automatically accessing and integrating large datasets from dissimilar resources due to representation, naming, and format conflicts, often requiring extensive processing and memory, and may pass errors uncorrected through sequential integration steps or necessitate consolidating data into a single repository.

Innovation Solution

A unified data integration system that detects and analyzes datasets across different formats and infrastructures, captures schema and data-element relationships, and virtually links records from various sources using fuzzy logic and metadata dictionaries, enabling seamless exploration and collaboration among analysts.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If data from multiple dissimilar sources is integrated into a single repository, then data accessibility and integration are improved, but processing time, memory requirements, and system complexity increase significantly

Engineering Contradiction:
Improvedata integration capabilityVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system segments the data integration process into independent resolution steps (representation conflict resolution, naming conflict resolution, format conflict resolution) that can be processed separately and in parallel, reducing the complexity of handling diverse data sources simultaneously

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system introduces a virtual warehouse as an intermediary layer between dissimilar data sources and analysis tools. This virtual warehouse resolves conflicts and standardizes data representations without requiring physical consolidation into a single repository, thereby reducing system complexity while maintaining integration capabilities

Inventive Principle:
Principle #24Intermediary (Mediator)

2Ease of manufacture

If sequential integration steps are used to resolve conflicts, then processing can be performed step-by-step, but errors pass on uncorrected from one step to the next

Engineering Contradiction:
Improveprocessing feasibilityVSAvoiderror correction
Core Design Contradiction:
Ease of manufactureVSReliability

Solution Approach 1:

The system implements feedback mechanisms at each resolution step where unresolved conflicts or errors are identified and returned to previous steps for reprocessing. This allows corrections to propagate backward through the integration process, ensuring that errors are not merely passed on but actually corrected before final data integration is complete

Inventive Principle:
Principle #23Feedback

3Quantity of substance

If extensive processing and scaling are performed to consolidate data, then data integration completeness is improved, but time required for integration increases

Engineering Contradiction:
Improvedata completenessVSAvoidintegration time
Core Design Contradiction:
Quantity of substanceVSLoss of time

Solution Approach 1:

The system performs preliminary actions by pre-resolving representation conflicts, naming conflicts, and format conflicts before the actual data consolidation occurs. By preparing conflict resolution mappings in advance, the system reduces the processing time required during data integration while maintaining completeness

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system uses partial action by resolving only the necessary conflicts required for a given query or analysis task rather than resolving all possible conflicts in the entire data warehouse. This selective approach maintains data completeness for relevant purposes while significantly reducing integration time

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS10127292B2Knowledge catalysts
Publication Date: 2018.11.13 UT BATTELLE LLC
  • US10127292B2 patent drawing
  • US10127292B2 patent drawing
  • US10127292B2 patent drawing

AI summary

A computer implemented method integrates data from remote disparate data sources by processing a non-transitory media. The non-transitory media stores instructions for detecting data sets in different formats hosted in a plurality of heterogeneous databases that are accessible through a distributed network. The method extracts schema data from the plurality of heterogeneous databases and identifies related fields in two or more of the heterogeneous databases. The method links the related fields in the two or more of the plurality of heterogeneous databases and makes the data accessible through a virtual warehouse. As schemas change, as new data sources and analysis artifacts are created, the computer implemented method and system can act as a meta-data store, a provenance tracking device, and/or a knowledge management service.