Ontology-Aware Active Learning for Schema Alignment
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing active learning techniques for graph neural network-based semantic schema alignment are limited in utilizing rich semantic information for effective and efficient sample selection and scaling to larger datasets, requiring extensive labeled training data and manual effort.
Innovation Solution
The proposed active learning framework (ALFA) exploits schema element properties and relationships to drive ontology-aware sample selection and label propagation, incorporating a novel ontology-aware sample selection algorithm, label propagation algorithm, and semantic blocking technique to minimize human labeling cost and scale to large datasets.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If existing active learning techniques are used for graph neural network-based semantic schema alignment, then the system can perform sample selection, but it requires extensive labeled training data and manual effort, and cannot effectively utilize rich semantic information for scaling to larger datasets
Solution Approach 1:
The patent introduces an ontology-aware sample selection algorithm that acts as an intermediary between the graph neural network and the training data. This algorithm utilizes rich semantic information from ontologies to intelligently select representative samples, reducing the need for extensive manually labeled training data while improving the system's ability to scale to larger datasets
Solution Approach 2:
The patent changes the parameters of sample selection by incorporating ontology-aware features and semantic information into the sampling criteria. This allows the system to adaptively select samples based on semantic similarity and ontology structure rather than random or uniform sampling, thereby reducing labeled data requirements while maintaining performance
2Measurement precision
If more labeled training data is used to improve model accuracy, then schema matching quality improves, but human labeling cost and manual effort increase substantially
Solution Approach 1:
The patent implements a self-service mechanism where the system automatically selects informative samples using ontology-aware algorithms and actively queries for labels only when necessary. This reduces human labeling cost by allowing the system to self-manage the training data selection process, leveraging semantic information to identify the most valuable samples for manual annotation
Solution Approach 2:
The patent incorporates feedback loops where the graph neural network model is iteratively trained and evaluated, with ontology-aware sample selection adjusting based on model performance. This feedback mechanism ensures that labeling efforts are focused on samples that most improve schema matching quality, reducing overall human labeling cost while maintaining high accuracy
3Quantity of substance
If existing active learning techniques are applied to larger datasets, then coverage increases, but the techniques cannot effectively scale due to limited semantic information utilization
Solution Approach 1:
The patent segments the large dataset into ontology-based clusters or groups, allowing the graph neural network to process and select samples from each segment using semantic information. This segmentation approach enables effective scaling to larger datasets by breaking down the complexity and utilizing ontology structure to guide sample selection across different data segments
Solution Approach 2:
The patent adds the dimension of ontology-aware semantic information to the sample selection process. By incorporating ontology hierarchies, semantic similarities, and conceptual relationships as additional dimensions for sample selection, the system can effectively scale to larger datasets while maintaining productivity through intelligent, semantics-driven sampling strategies
Data Source
AI summary
Embodiments are related to a technique for active learning for graph neural network based semantic schema alignment. The technique includes generating, by a first machine learning model executed on a processor, node embeddings having node pairs of a first schema and a second schema. The technique includes predicting, by a second machine learning model executed on the processor, a label output for the node pairs. The technique includes clustering the node pairs into a cluster output, determining that the label output and the cluster output are in a disagreement for at least one node pair of the node pairs, and in response to displaying the at least one node pair to a subject matter expert to generate a label for the at least one node pair, using the label for the at least one node pair as training data to further train the second machine learning model.


