Ontology Matching Threshold Computation for Semantic Entity Identification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional methods for identifying semantically similar entities in text data are complex and ad-hoc, making it challenging to accurately compute thresholds for ontology matching, which is crucial for data privacy and security in identifying sensitive information.

Innovation Solution

A method and system that generate an Entity Relationship (ER) model from text descriptions, convert it into ontologies, and use ontology matching algorithms to compute confidence values for alignments, with a pre-determined threshold calculated based on symmetric and transitive properties to select semantically similar entities.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Extent of automation

If conventional automated threshold computation methods (k-means clustering and fuzzy logic) are used, then the threshold can be computed automatically, but the methods become complex and ad-hoc in nature

Engineering Contradiction:
Improveautomatic threshold computationVSAvoidmethod complexity
Core Design Contradiction:
Extent of automationVSDevice complexity

Solution Approach 1:

The patent changes the parameter approach from manual/semi-automatic threshold setting to an automated statistical parameter computation. The threshold is derived automatically from the data distribution using percentile-based methods, transforming the complex ad-hoc parameter selection into a systematic statistical process that reduces manual intervention while maintaining clarity in the computation methodology.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The system performs self-service by automatically computing the threshold based on the statistical properties of the ontology alignment scores themselves. The threshold is derived from the data distribution (using percentiles) without requiring external manual input or complex pre-computed parameters, allowing the system to self-calibrate the threshold based on the actual alignment results.

Inventive Principle:
Principle #25Self-service

2Extent of automation

If conventional automatic threshold computation methods are used, then the threshold can be computed without manual intervention, but the methods are forced fit for the ontology matching task

Engineering Contradiction:
Improveautomatic threshold computationVSAvoidadaptability to ontology matching
Core Design Contradiction:
Extent of automationVSAdaptability or versatility

Solution Approach 1:

The patent adapts the threshold computation to the specific needs of ontology matching by using percentile-based statistical parameters that can be adjusted according to different matching scenarios. The threshold is computed as a percentile of the alignment scores, allowing flexible adaptation to different ontology domains and matching requirements without forcing a single fixed methodology.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The threshold computation is made dynamic by calculating it based on the actual distribution of ontology alignment scores in the dataset. Rather than using a static pre-defined threshold, the system dynamically computes the threshold percentile based on the specific ontology matching task at hand, allowing the threshold to adapt to different data characteristics and matching scenarios.

Inventive Principle:
Principle #15Dynamics

3Reliability

If a threshold is applied to identify semantically similar entities, then sensitive data can be identified, but the accuracy of identification becomes challenging with conventional methods

Engineering Contradiction:
Improvesensitive data identificationVSAvoidsemantic similarity identification accuracy
Core Design Contradiction:
ReliabilityVSMeasurement precision

Solution Approach 1:

The system incorporates feedback mechanisms where the threshold computation is based on the actual ontology alignment scores obtained from the text data. The percentile-based threshold is dynamically adjusted according to the distribution of alignment scores, creating a feedback loop that continuously refines the threshold to optimize both sensitivity and precision in identifying semantically similar entities and sensitive data.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent replaces the mechanical/manual threshold setting process with a statistical computational approach. Instead of relying on manual calibration or fixed mechanical thresholds, the system uses statistical percentiles and data-driven computations to determine the threshold, thereby improving measurement precision and accuracy in identifying semantic similarities without requiring manual intervention.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS11762885B2System and method for identifying semantic similarity
Publication Date: 2023.09.19 TATA CONSULTANCY SERVICES LTD
  • US11762885B2 patent drawing
  • US11762885B2 patent drawing
  • US11762885B2 patent drawing

AI summary

Protecting consumer data is a key responsibility of an organization. The method of protection need to compare the data with policy documents. Ontology and a threshold to select an optimal match plays a key role in such comparison. The conventional automatic threshold computation methods are complex and not based on semantic similarity. The present disclosure generates an Entity Relationship (ER) model from an input document and is converted into a first ontology. The first ontology and a second ontology obtained from a relational database are compared by an ontology matching algorithm. Further, the plurality of many to many correspondences are optimized to one to one correspondence by an optimization method. Further, a plurality of optimal one to one correspondence is generated based on a threshold. The threshold is computed based on symmetric and transitive property. Further, semantically similar entities are selected based on the optimal one to one correspondence.