Ontology Matching Threshold Computation for Semantic Entity Identification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional methods for identifying semantically similar entities in text data are complex and ad-hoc, making it challenging to accurately compute thresholds for ontology matching, which is crucial for data privacy and security in identifying sensitive information.
Innovation Solution
A method and system that generate an Entity Relationship (ER) model from text descriptions, convert it into ontologies, and use ontology matching algorithms to compute confidence values for alignments, with a pre-determined threshold calculated based on symmetric and transitive properties to select semantically similar entities.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Extent of automation
If conventional automated threshold computation methods (k-means clustering and fuzzy logic) are used, then the threshold can be computed automatically, but the methods become complex and ad-hoc in nature
Solution Approach 1:
The patent changes the parameter approach from manual/semi-automatic threshold setting to an automated statistical parameter computation. The threshold is derived automatically from the data distribution using percentile-based methods, transforming the complex ad-hoc parameter selection into a systematic statistical process that reduces manual intervention while maintaining clarity in the computation methodology.
Solution Approach 2:
The system performs self-service by automatically computing the threshold based on the statistical properties of the ontology alignment scores themselves. The threshold is derived from the data distribution (using percentiles) without requiring external manual input or complex pre-computed parameters, allowing the system to self-calibrate the threshold based on the actual alignment results.
2Extent of automation
If conventional automatic threshold computation methods are used, then the threshold can be computed without manual intervention, but the methods are forced fit for the ontology matching task
Solution Approach 1:
The patent adapts the threshold computation to the specific needs of ontology matching by using percentile-based statistical parameters that can be adjusted according to different matching scenarios. The threshold is computed as a percentile of the alignment scores, allowing flexible adaptation to different ontology domains and matching requirements without forcing a single fixed methodology.
Solution Approach 2:
The threshold computation is made dynamic by calculating it based on the actual distribution of ontology alignment scores in the dataset. Rather than using a static pre-defined threshold, the system dynamically computes the threshold percentile based on the specific ontology matching task at hand, allowing the threshold to adapt to different data characteristics and matching scenarios.
3Reliability
If a threshold is applied to identify semantically similar entities, then sensitive data can be identified, but the accuracy of identification becomes challenging with conventional methods
Solution Approach 1:
The system incorporates feedback mechanisms where the threshold computation is based on the actual ontology alignment scores obtained from the text data. The percentile-based threshold is dynamically adjusted according to the distribution of alignment scores, creating a feedback loop that continuously refines the threshold to optimize both sensitivity and precision in identifying semantically similar entities and sensitive data.
Solution Approach 2:
The patent replaces the mechanical/manual threshold setting process with a statistical computational approach. Instead of relying on manual calibration or fixed mechanical thresholds, the system uses statistical percentiles and data-driven computations to determine the threshold, thereby improving measurement precision and accuracy in identifying semantic similarities without requiring manual intervention.
Data Source
AI summary
Protecting consumer data is a key responsibility of an organization. The method of protection need to compare the data with policy documents. Ontology and a threshold to select an optimal match plays a key role in such comparison. The conventional automatic threshold computation methods are complex and not based on semantic similarity. The present disclosure generates an Entity Relationship (ER) model from an input document and is converted into a first ontology. The first ontology and a second ontology obtained from a relational database are compared by an ontology matching algorithm. Further, the plurality of many to many correspondences are optimized to one to one correspondence by an optimization method. Further, a plurality of optimal one to one correspondence is generated based on a threshold. The threshold is computed based on symmetric and transitive property. Further, semantically similar entities are selected based on the optimal one to one correspondence.


