Concept Set Similarity Analysis Using Hierarchical Path Metrics

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional search engines rely on text-based searching, which limits the ability to find information by matching text, failing to effectively determine similarities between concept sets, especially in semantically enriched web pages.

Innovation Solution

A system that includes a concept analysis engine with a taxonomy manager, concept pair engine, hierarchical path engine, and concept similarity engine to determine concept pairs and their similarity values based on nondiverging intersections and weighted sums, enhancing the identification of similar concept sets.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If text-based searching is used, then search simplicity is maintained, but search capability and information retrieval effectiveness deteriorate

Engineering Contradiction:
Improvesearch simplicityVSAvoidsearch capability
Core Design Contradiction:
Ease of operationVSProductivity

Solution Approach 1:

The patent introduces concept sets as an intermediary layer between text-based queries and information retrieval. Instead of directly searching text, the system converts queries into concept sets, compares them with concept sets associated with web pages, and retrieves relevant pages based on concept set similarity. This mediator approach maintains user-friendly operation while significantly improving search capability through semantic understanding.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Device complexity

If conventional text-based search is used, then system complexity is low, but measurement precision of concept similarity deteriorates

Engineering Contradiction:
Improvesystem complexityVSAvoidconcept similarity determination
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent transitions from one-dimensional text matching to multi-dimensional concept set comparison. By representing both queries and web pages as sets of concepts with hierarchical relationships, the system evaluates similarity across multiple dimensions including concept presence, hierarchical depth, and set overlap, thereby achieving precise concept similarity determination while managing complexity through structured representation.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Loss of information

If semantic information enrichment is added to web pages, then information accessibility improves, but data processing complexity increases

Engineering Contradiction:
Improveinformation accessibilityVSAvoiddata processing complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent segments semantic information into discrete concept sets that can be independently processed and compared. Each web page is associated with a structured set of concepts rather than processing entire pages or complex metadata, allowing efficient retrieval and comparison operations while maintaining comprehensive semantic information accessibility.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS7860855B2Method and system for analyzing similarity of concept sets
Publication Date: 2010.12.28 SAP SE
  • US7860855B2 patent drawing
  • US7860855B2 patent drawing
  • US7860855B2 patent drawing

AI summary

A method and system are described for determining similar concept sets. An example method includes obtaining taxonomies, each including one root node and hierarchically ordered paths; receiving first and second sets each including set concepts; determining concept pairs, each including a first and second set concept; determining lengths of nondiverging intersections of first and second subpaths from the root node to first and second concept nodes, and associated lengths of first and second portions of the subpaths from a last concept node included in the nondiverging intersection to the first and second concept nodes; determining pairwise similarity values based on ratios based on associated lengths of nondiverging intersections and the associated lengths of the first and second portions; and determining a concept set similarity value based on a weighted sum of the pairwise similarity values associated with optimal selected ones of the concept pairs.