Ontology Mapping for Digital Twin Data Homogenization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing data integration methods for machine learning models require significant human intervention and effort to homogenize datasets with different ontologies, leading to scalability challenges in creating high-fidelity digital twins and suboptimal performance.

Innovation Solution

A computer-implemented method that automatically maps and scores concepts between different ontologies, merges them based on performance relevance, and generates a homogenized dataset for machine learning models, using techniques like Pearson correlation and Large Language Models (LLMs) to optimize model performance.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual ontology mapping and data homogenization is performed, then data integration quality is improved, but human effort and time consumption increase significantly

Engineering Contradiction:
Improvedata integration qualityVSAvoidhuman effort and time consumption
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs self-service by automatically mapping ontologies and homogenizing datasets through machine learning algorithms. The ontology mapper module autonomously identifies and maps concepts between different ontologies, while the data homogenizer module automatically transforms datasets, eliminating the need for manual human intervention in these tasks.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces manual mechanical processes with automated computational systems. Manual ontology mapping is substituted by an automated ontology mapper that uses machine learning, and manual data homogenization is replaced by an automated data homogenizer that applies transformation rules and algorithms to align datasets.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Quantity of substance

If comprehensive ontology mapping is performed to include all concepts, then data completeness is improved, but computational complexity and processing time increase

Engineering Contradiction:
Improvedata completenessVSAvoidcomputational complexity
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The system extracts and filters only the most relevant concepts and features from comprehensive ontologies during the mapping process. The machine learning algorithms identify and select key concepts that are most useful for the specific task, excluding redundant or less important concepts, thereby reducing computational complexity while maintaining data completeness for the intended application.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent applies local quality by tailoring the ontology mapping and data homogenization process to specific local requirements and contexts. The system adapts the mapping strategy based on the specific datasets and application domain, focusing computational resources on the most critical concepts and relationships rather than uniformly processing all concepts across all ontologies.

Inventive Principle:
Principle #3Local quality

3Reliability

If data homogenization is performed to standardize datasets, then machine learning model performance is improved, but data transformation complexity increases

Engineering Contradiction:
Improvemachine learning model performanceVSAvoiddata transformation complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system performs preliminary action by pre-defining transformation rules and templates for data homogenization. The ontology mapper pre-establishes mapping relationships between different ontologies, and the data homogenizer uses these pre-defined rules to automatically transform datasets, reducing the complexity of data transformation while ensuring consistent standardization that improves machine learning model performance.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20250348471A1Machine learning-driven data integration for data spaces and digital twins
Publication Date: 2025.11.13 NEC LAB EURO GMBH
  • US20250348471A1 patent drawing
  • US20250348471A1 patent drawing
  • US20250348471A1 patent drawing

AI summary

A computer-implemented method for homogenizing datasets that have different ontologies to optimize performance of machine learning models that use the datasets includes mapping concepts between the different ontologies of the datasets based on ontology matching. The concepts are scored based on a relation between the concepts to identify certain ones of the concepts that are more important to improving the performance of the machine learning models. The different ontologies are merged based on the scoring to generate a merged ontology that includes the identified concepts. The datasets are transformed into a homogenized dataset according to the merged ontology. A machine learning model is generated based on the homogenized dataset. The method has applications including, but not limited to, use cases in computational biology, medical AI and healthcare, cyberthreat security, public safety and smart cities for optimizing machine learning processes or supporting decision making.