Ontology-Aware Active Learning for Schema Alignment

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing active learning techniques for graph neural network-based semantic schema alignment are limited in utilizing rich semantic information for effective and efficient sample selection and scaling to larger datasets, requiring extensive labeled training data and manual effort.

Innovation Solution

The proposed active learning framework (ALFA) exploits schema element properties and relationships to drive ontology-aware sample selection and label propagation, incorporating a novel ontology-aware sample selection algorithm, label propagation algorithm, and semantic blocking technique to minimize human labeling cost and scale to large datasets.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If existing active learning techniques are used for graph neural network-based semantic schema alignment, then the system can perform sample selection, but it requires extensive labeled training data and manual effort, and cannot effectively utilize rich semantic information for scaling to larger datasets

Engineering Contradiction:
Improveutilization of semantic informationVSAvoidlabeled training data
Core Design Contradiction:
Adaptability or versatilityVSQuantity of substance

Solution Approach 1:

The patent introduces an ontology-aware sample selection algorithm that acts as an intermediary between the graph neural network and the training data. This algorithm utilizes rich semantic information from ontologies to intelligently select representative samples, reducing the need for extensive manually labeled training data while improving the system's ability to scale to larger datasets

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent changes the parameters of sample selection by incorporating ontology-aware features and semantic information into the sampling criteria. This allows the system to adaptively select samples based on semantic similarity and ontology structure rather than random or uniform sampling, thereby reducing labeled data requirements while maintaining performance

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If more labeled training data is used to improve model accuracy, then schema matching quality improves, but human labeling cost and manual effort increase substantially

Engineering Contradiction:
Improveschema matching qualityVSAvoidhuman labeling cost
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent implements a self-service mechanism where the system automatically selects informative samples using ontology-aware algorithms and actively queries for labels only when necessary. This reduces human labeling cost by allowing the system to self-manage the training data selection process, leveraging semantic information to identify the most valuable samples for manual annotation

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent incorporates feedback loops where the graph neural network model is iteratively trained and evaluated, with ontology-aware sample selection adjusting based on model performance. This feedback mechanism ensures that labeling efforts are focused on samples that most improve schema matching quality, reducing overall human labeling cost while maintaining high accuracy

Inventive Principle:
Principle #23Feedback

3Quantity of substance

If existing active learning techniques are applied to larger datasets, then coverage increases, but the techniques cannot effectively scale due to limited semantic information utilization

Engineering Contradiction:
Improvedataset sizeVSAvoidscaling efficiency
Core Design Contradiction:
Quantity of substanceVSProductivity

Solution Approach 1:

The patent segments the large dataset into ontology-based clusters or groups, allowing the graph neural network to process and select samples from each segment using semantic information. This segmentation approach enables effective scaling to larger datasets by breaking down the complexity and utilizing ontology structure to guide sample selection across different data segments

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent adds the dimension of ontology-aware semantic information to the sample selection process. By incorporating ontology hierarchies, semantic similarities, and conceptual relationships as additional dimensions for sample selection, the system can effectively scale to larger datasets while maintaining productivity through intelligent, semantics-driven sampling strategies

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS20240330693A1Active learning for graph neural network based semantic schema alignment
Publication Date: 2024.10.03 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US20240330693A1 patent drawing
  • US20240330693A1 patent drawing
  • US20240330693A1 patent drawing

AI summary

Embodiments are related to a technique for active learning for graph neural network based semantic schema alignment. The technique includes generating, by a first machine learning model executed on a processor, node embeddings having node pairs of a first schema and a second schema. The technique includes predicting, by a second machine learning model executed on the processor, a label output for the node pairs. The technique includes clustering the node pairs into a cluster output, determining that the label output and the cluster output are in a disagreement for at least one node pair of the node pairs, and in response to displaying the at least one node pair to a subject matter expert to generate a label for the at least one node pair, using the label for the at least one node pair as training data to further train the second machine learning model.