Semantic Data Lake for Automatic Heterogeneous Data Integration

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current data analytics pipelines are inflexible and require manual effort to integrate new data sources, hindering the generation of actionable insights due to the lack of automatic evolution and integration of heterogeneous data sources, which is challenging for data scientists and subject-matter experts without a data-science background.

Innovation Solution

A system and method that utilize a data integration module to create semantic annotations and a knowledge graph, enabling the integration of raw data into a semantic data lake, making the data interpretable and allowing for automatic evolution with new data sources, facilitated by a knowledge reasoning engine that supports interactive workflow creation and modification based on capabilities rather than raw data bits.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If manual integration of data sources is performed by data scientists, then data integration capability is achieved, but time consumption and complexity increase significantly

Engineering Contradiction:
Improvedata integration capabilityVSAvoidtime consumption
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The patent introduces a semantic data lake as an intermediary layer between raw data sources and analytics workflows. This semantic layer automatically creates standardized representations of data from diverse sources, eliminating the need for manual data scientist intervention. The semantic data lake acts as a mediator that harmonizes data from multiple sources using predefined schemas and ontologies, thereby reducing integration time while maintaining adaptability.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system performs preliminary actions by pre-defining data schemas, ontologies, and semantic models before data integration occurs. These pre-configured frameworks enable automatic data classification, validation, and harmonization when new data sources are connected. This preliminary preparation eliminates the need for time-consuming manual integration efforts for each new data source.

Inventive Principle:
Principle #10Preliminary action

2Manufacturing precision

If data scientists perform bit-level integration of data sources, then precise data integration is achieved, but the process becomes slow and requires specialized expertise

Engineering Contradiction:
Improvedata integration precisionVSAvoidintegration speed
Core Design Contradiction:
Manufacturing precisionVSProductivity

Solution Approach 1:

The patent replaces the mechanical manual process of bit-level data integration with an automated semantic processing system. Instead of data scientists manually manipulating data bits, the system uses semantic annotations, ontologies, and automated reasoning engines to interpret and integrate data. This substitution maintains integration precision through semantic validation while dramatically increasing productivity through automation.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system changes the parameter of data representation from raw bits to semantic annotations. By transforming data into annotated formats with embedded meaning and context, the system enables automated processing that maintains precision through semantic rules while improving speed. The parameter change from binary data to semantic data structures allows for more efficient automated integration.

Inventive Principle:
Principle #35Parameter changes

3Stability of the object's composition

If the existing data pipeline is not evolved automatically, then system stability is maintained, but new data sources cannot be integrated quickly

Engineering Contradiction:
Improvesystem stabilityVSAvoidability to integrate new data sources
Core Design Contradiction:
Stability of the object's compositionVSAdaptability or versatility

Solution Approach 1:

The patent introduces dynamic capabilities to the data pipeline through the semantic data lake, which can automatically adapt its schema and ontologies when new data sources are connected. The system maintains stability through its core semantic framework while dynamically evolving to accommodate new data types and sources. This dynamic adaptation allows the pipeline to remain stable in its core functionality while being versatile in integrating new sources.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The semantic data lake provides a universal interface that can handle multiple types of data sources through a common semantic framework. This multi-functional approach allows the same infrastructure to process diverse data types (structured, unstructured, semi-structured) while maintaining system stability. The universal semantic layer acts as an adapter that preserves core system stability while enabling versatility in data source integration.

Inventive Principle:
Principle #6Universality (Multi-functionality)

4Reliability

If subject-matter experts collaborate with data scientists for data integration, then domain knowledge is incorporated, but the collaboration process is time-consuming

Engineering Contradiction:
Improvedomain knowledge accuracyVSAvoidcollaboration time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system enables self-service by allowing subject-matter experts to define their own data sources and schemas using the semantic framework without requiring data scientist involvement. The automated semantic processing system handles the technical integration tasks, while domain experts focus on providing semantic annotations and business rules. This self-service capability maintains domain knowledge accuracy while eliminating time-consuming collaboration cycles.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS20230351209A1Systems and methods for defining data analytics pipelines
Publication Date: 2023.11.02 SIEMENS MOBILITY INC
  • US20230351209A1 patent drawing
  • US20230351209A1 patent drawing
  • US20230351209A1 patent drawing

AI summary

A system for defining data analytics pipelines with a processor and a memory includes a data source with raw data, a semantic data lake, and a data integration module, wherein the data integration module is configured via computer executable instructions to create semantic annotations that describe a capability and a structure of the raw data of the data source, create or modify a knowledge graph utilizing the semantic annotations, and integrate the raw data and the semantic annotations into the semantic data lake, wherein the raw data are interpretable via the knowledge graph and the semantic annotations.