Ontology Mapping for Digital Twin Data Homogenization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data integration methods for machine learning models require significant human intervention and effort to homogenize datasets with different ontologies, leading to scalability challenges in creating high-fidelity digital twins and suboptimal performance.
Innovation Solution
A computer-implemented method that automatically maps and scores concepts between different ontologies, merges them based on performance relevance, and generates a homogenized dataset for machine learning models, using techniques like Pearson correlation and Large Language Models (LLMs) to optimize model performance.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual ontology mapping and data homogenization is performed, then data integration quality is improved, but human effort and time consumption increase significantly
Solution Approach 1:
The system performs self-service by automatically mapping ontologies and homogenizing datasets through machine learning algorithms. The ontology mapper module autonomously identifies and maps concepts between different ontologies, while the data homogenizer module automatically transforms datasets, eliminating the need for manual human intervention in these tasks.
Solution Approach 2:
The patent replaces manual mechanical processes with automated computational systems. Manual ontology mapping is substituted by an automated ontology mapper that uses machine learning, and manual data homogenization is replaced by an automated data homogenizer that applies transformation rules and algorithms to align datasets.
2Quantity of substance
If comprehensive ontology mapping is performed to include all concepts, then data completeness is improved, but computational complexity and processing time increase
Solution Approach 1:
The system extracts and filters only the most relevant concepts and features from comprehensive ontologies during the mapping process. The machine learning algorithms identify and select key concepts that are most useful for the specific task, excluding redundant or less important concepts, thereby reducing computational complexity while maintaining data completeness for the intended application.
Solution Approach 2:
The patent applies local quality by tailoring the ontology mapping and data homogenization process to specific local requirements and contexts. The system adapts the mapping strategy based on the specific datasets and application domain, focusing computational resources on the most critical concepts and relationships rather than uniformly processing all concepts across all ontologies.
3Reliability
If data homogenization is performed to standardize datasets, then machine learning model performance is improved, but data transformation complexity increases
Solution Approach 1:
The system performs preliminary action by pre-defining transformation rules and templates for data homogenization. The ontology mapper pre-establishes mapping relationships between different ontologies, and the data homogenizer uses these pre-defined rules to automatically transform datasets, reducing the complexity of data transformation while ensuring consistent standardization that improves machine learning model performance.
Data Source
AI summary
A computer-implemented method for homogenizing datasets that have different ontologies to optimize performance of machine learning models that use the datasets includes mapping concepts between the different ontologies of the datasets based on ontology matching. The concepts are scored based on a relation between the concepts to identify certain ones of the concepts that are more important to improving the performance of the machine learning models. The different ontologies are merged based on the scoring to generate a merged ontology that includes the identified concepts. The datasets are transformed into a homogenized dataset according to the merged ontology. A machine learning model is generated based on the homogenized dataset. The method has applications including, but not limited to, use cases in computational biology, medical AI and healthcare, cyberthreat security, public safety and smart cities for optimizing machine learning processes or supporting decision making.


