Text Analysis Result Merging via Corrected Jaccard Factors

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Merging results from multiple text mining services is challenging due to differences in taxonomies and weak metadata, leading to insufficient ontology matching results, as existing approaches struggle to identify equal, hierarchical, and associative mappings between taxonomies.

Innovation Solution

A system that enriches service taxonomies with instance information using a novel instance-based matching technique and metric, allowing for the automatic identification of equal, hierarchical, and associative mappings by calculating corrected Jaccard factors to merge results from multiple text analysis services.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If multiple text mining services are used to extract information, then the quantity and diversity of extracted information increases, but the difficulty of merging results increases due to different taxonomies

Engineering Contradiction:
Improvequantity of extracted informationVSAvoidcomplexity of merging results
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The patent introduces an intermediary component (taxonomy mapping system) that mediates between different text mining services' taxonomies. This intermediary automatically computes mappings between disparate taxonomies using ontology matching, enabling results from multiple services to be merged without manual intervention. The intermediary resolves the contradiction by providing an automated bridge that handles the complexity of taxonomy alignment.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent changes the parameters of the merging process by transforming the problem from manual taxonomy alignment to automated instance-based matching. By computing similarity coefficients (such as Jaccard index) between taxonomy elements based on instance data, the system dynamically adjusts the merging process to handle different taxonomy structures, thereby reducing the complexity of integrating results from multiple services.

Inventive Principle:
Principle #35Parameter changes

2Extent of automation

If ontology matching systems are used to compute mappings between taxonomies, then the automation level increases, but the quality of matching results decreases due to weak metadata

Engineering Contradiction:
Improveautomation of taxonomy mappingVSAvoidprecision of mapping results
Core Design Contradiction:
Extent of automationVSMeasurement precision

Solution Approach 1:

The patent applies preliminary action by enriching taxonomy elements with instance data before performing matching. Instead of directly matching taxonomy elements with weak metadata, the system first populates them with concrete instances from text mining results. This preliminary enrichment provides additional information that improves the precision of subsequent automated matching operations.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent substitutes traditional metadata-based matching mechanisms with instance-based similarity computation. Instead of relying on weak textual metadata to compute mappings, the system uses concrete instances associated with taxonomy elements to calculate similarity coefficients. This substitution replaces the inadequate mechanical matching process with a more robust instance-driven approach, improving mapping precision while maintaining automation.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Ease of manufacture

If traditional Jaccard factor is used for matching instances, then the calculation simplicity is maintained, but the accuracy of identifying taxonomy relationships decreases

Engineering Contradiction:
Improvesimplicity of calculationVSAvoidaccuracy of relationship identification
Core Design Contradiction:
Ease of manufactureVSMeasurement precision

Solution Approach 1:

The patent modifies the Jaccard factor by introducing weighting parameters that reflect the importance of different instance types and matching criteria. By changing the parameters of the similarity calculation (adding weights for different factors such as instance frequency, type hierarchy, and co-occurrence patterns), the system maintains the computational framework of the Jaccard index while significantly improving the accuracy of taxonomy relationship identification.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS9047347B2System and method of merging text analysis results
Publication Date: 2015.06.02 SAP SE
  • US9047347B2 patent drawing
  • US9047347B2 patent drawing
  • US9047347B2 patent drawing

AI summary

A system and method of merging text analysis results. The system uses a set of three corrected, weakened Jaccard factors to determine whether the respective results of multiple text analysis operations are equal, subtypes of each other or associated with each other, in order to merge the results.