Machine Learning Ontology Mapping for Dataset Annotation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The annotation of datasets for training machine learning models is cumbersome and time-consuming for skilled personnel, influencing the performance and accuracy of the models, and there is a need for improving this process with minimal effort.

Innovation Solution

A system and method that uses a machine learning model to predict ontology labels for dataset columns, allowing user input to establish relations, train the model, and generate mappings between datasets and ontologies, thereby automating the annotation process.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If skilled personnel manually annotate datasets for machine learning model training, then the quality and accuracy of annotations can be maintained, but the process becomes cumbersome and time-consuming

Engineering Contradiction:
Improveannotation accuracyVSAvoidannotation time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system enables self-service annotation by allowing users to automatically generate ontology labels for dataset columns using a machine learning model, eliminating the need for manual annotation by skilled personnel while maintaining reasonable accuracy through model predictions

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the manual mechanical process of skilled personnel annotating data with an automated machine learning-based system that predicts ontology labels, substituting human effort with computational intelligence

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Reliability

If large datasets are annotated with appropriate labels by skilled personnel to train machine learning models, then the performance and accuracy of the models improve, but the annotation process becomes quite cumbersome and time consuming

Engineering Contradiction:
Improvemodel performanceVSAvoidannotation effort
Core Design Contradiction:
ReliabilityVSEase of manufacture

Solution Approach 1:

The system allows the machine learning model to train itself by using automatically generated ontology labels from the ontology mapping process, eliminating the need for separate manual annotation efforts while still providing sufficient training data

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system performs preliminary ontology label prediction and mapping before model training, preparing the annotated dataset automatically in advance, which then serves as training data without requiring additional manual annotation efforts

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If the amount and quality of annotations are increased to improve machine learning model accuracy, then the model performance improves, but the effort required from skilled personnel increases

Engineering Contradiction:
Improvemodel accuracyVSAvoidannotation process complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The ontology mapping system serves multiple functions simultaneously: it structures the dataset, generates ontology labels, creates mappings between data and ontology, and provides training data for machine learning models, all through a single automated process

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent replaces the complex manual process of skilled personnel analyzing and annotating data with an automated machine learning-based ontology mapping system that handles the complexity computationally

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS12461959B2System and method for data management
Publication Date: 2025.11.04 SIEMENS SCHWEIZ AG
  • US12461959B2 patent drawing
  • US12461959B2 patent drawing
  • US12461959B2 patent drawing

AI summary

A system and method for data management is provided. The method includes obtaining a dataset from a data source by a processing unit. The dataset includes a plurality of datapoints, and each of the datapoints belongs to a column among a plurality of columns. Further, an ontology label for at least one column in the dataset is predicted using a machine learning model. The predicted ontology label is associated with an ontology comprising a plurality of ontology labels. Further, a mapping between the dataset and the ontology is generated based on the relation between the predicted ontology label and the column. Furthermore, the datapoints are classified with respect to the ontology labels based on the mapping generated. The classified datasets are outputted on a user interface.