Word Cloud Generation Using Folksonomy Tag Boosting
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing tag-cloud generation techniques face difficulties in creating effective representations for items in sparsely tagged domains, as they rely on manual tags which are not always available or sufficient, leading to inferior automatic tag-clouds generated from statistically significant but less relevant terms.
Innovation Solution
The 'tag-boost' method extracts terms from content items using statistical selection criteria and weights them based on their probability of being used as tags, leveraging a folksonomy to promote frequently used terms for enhanced visual representation in word-clouds, even in non-tagged or poorly tagged domains.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If extracted terms are used as alternative tags in sparsely tagged domains, then tag-cloud generation is enabled, but the quality of tags deteriorates
Solution Approach 1:
The patent introduces a folksonomy-based weighting mechanism as an intermediary to bridge the gap between extracted terms and meaningful tags. The system extracts terms from content items and then applies a weighting function that incorporates folksonomy data to evaluate and rank these terms, selecting only those with sufficient weight to represent meaningful tags. This intermediary weighting layer transforms raw extracted terms into quality-filtered tag candidates.
Solution Approach 2:
The patent changes the parameter for tag selection from simple statistical frequency to a composite weight that incorporates folksonomy-based probability of tag usage. By transforming the selection criterion from raw term frequency to weighted score (combining term frequency with folksonomy probability), the system achieves both automation capability and tag quality in sparsely tagged domains.
2Extent of automation
If statistical selection criteria are used to extract terms, then automatic term extraction is achieved, but term relevance to content deteriorates
Solution Approach 1:
The patent implements a feedback mechanism where the folksonomy data (manual tag usage patterns) is used to evaluate and refine the automatically extracted terms. The weighting function incorporates feedback from folksonomy probability calculations, allowing the system to learn from historical tag usage patterns and adjust the relevance assessment of extracted terms accordingly. This feedback loop ensures that automatically extracted terms are filtered through the lens of actual user tagging behavior.
Data Source
AI summary
Method, system, and computer program product for automatic generation of a word-cloud for a content item are provided. The method includes: extracting terms from a content item using statistical selection criteria; weighting a term by a probability that the term is used as a tag; and generating a visual representation of terms with enhanced representation of terms according to the weighting. Weighting a term by a probability that the term is used as a tag may include determining the relative frequency of the term in a folksonomy of tag terms for a domain.


