Concept Set Similarity Analysis Using Hierarchical Path Metrics
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional search engines rely on text-based searching, which limits the ability to find information by matching text, failing to effectively determine similarities between concept sets, especially in semantically enriched web pages.
Innovation Solution
A system that includes a concept analysis engine with a taxonomy manager, concept pair engine, hierarchical path engine, and concept similarity engine to determine concept pairs and their similarity values based on nondiverging intersections and weighted sums, enhancing the identification of similar concept sets.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If text-based searching is used, then search simplicity is maintained, but search capability and information retrieval effectiveness deteriorate
Solution Approach 1:
The patent introduces concept sets as an intermediary layer between text-based queries and information retrieval. Instead of directly searching text, the system converts queries into concept sets, compares them with concept sets associated with web pages, and retrieves relevant pages based on concept set similarity. This mediator approach maintains user-friendly operation while significantly improving search capability through semantic understanding.
2Device complexity
If conventional text-based search is used, then system complexity is low, but measurement precision of concept similarity deteriorates
Solution Approach 1:
The patent transitions from one-dimensional text matching to multi-dimensional concept set comparison. By representing both queries and web pages as sets of concepts with hierarchical relationships, the system evaluates similarity across multiple dimensions including concept presence, hierarchical depth, and set overlap, thereby achieving precise concept similarity determination while managing complexity through structured representation.
3Loss of information
If semantic information enrichment is added to web pages, then information accessibility improves, but data processing complexity increases
Solution Approach 1:
The patent segments semantic information into discrete concept sets that can be independently processed and compared. Each web page is associated with a structured set of concepts rather than processing entire pages or complex metadata, allowing efficient retrieval and comparison operations while maintaining comprehensive semantic information accessibility.
Data Source
AI summary
A method and system are described for determining similar concept sets. An example method includes obtaining taxonomies, each including one root node and hierarchically ordered paths; receiving first and second sets each including set concepts; determining concept pairs, each including a first and second set concept; determining lengths of nondiverging intersections of first and second subpaths from the root node to first and second concept nodes, and associated lengths of first and second portions of the subpaths from a last concept node included in the nondiverging intersection to the first and second concept nodes; determining pairwise similarity values based on ratios based on associated lengths of nondiverging intersections and the associated lengths of the first and second portions; and determining a concept set similarity value based on a weighted sum of the pairwise similarity values associated with optimal selected ones of the concept pairs.


