Word Cloud Generation Using Folksonomy Tag Boosting

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing tag-cloud generation techniques face difficulties in creating effective representations for items in sparsely tagged domains, as they rely on manual tags which are not always available or sufficient, leading to inferior automatic tag-clouds generated from statistically significant but less relevant terms.

Innovation Solution

The 'tag-boost' method extracts terms from content items using statistical selection criteria and weights them based on their probability of being used as tags, leveraging a folksonomy to promote frequently used terms for enhanced visual representation in word-clouds, even in non-tagged or poorly tagged domains.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If extracted terms are used as alternative tags in sparsely tagged domains, then tag-cloud generation is enabled, but the quality of tags deteriorates

Engineering Contradiction:
Improvetag-cloud generation capabilityVSAvoidtag quality
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent introduces a folksonomy-based weighting mechanism as an intermediary to bridge the gap between extracted terms and meaningful tags. The system extracts terms from content items and then applies a weighting function that incorporates folksonomy data to evaluate and rank these terms, selecting only those with sufficient weight to represent meaningful tags. This intermediary weighting layer transforms raw extracted terms into quality-filtered tag candidates.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent changes the parameter for tag selection from simple statistical frequency to a composite weight that incorporates folksonomy-based probability of tag usage. By transforming the selection criterion from raw term frequency to weighted score (combining term frequency with folksonomy probability), the system achieves both automation capability and tag quality in sparsely tagged domains.

Inventive Principle:
Principle #35Parameter changes

2Extent of automation

If statistical selection criteria are used to extract terms, then automatic term extraction is achieved, but term relevance to content deteriorates

Engineering Contradiction:
Improveautomatic term extractionVSAvoidterm relevance
Core Design Contradiction:
Extent of automationVSReliability

Solution Approach 1:

The patent implements a feedback mechanism where the folksonomy data (manual tag usage patterns) is used to evaluate and refine the automatically extracted terms. The weighting function incorporates feedback from folksonomy probability calculations, allowing the system to learn from historical tag usage patterns and adjust the relevance assessment of extracted terms accordingly. This feedback loop ensures that automatically extracted terms are filtered through the lens of actual user tagging behavior.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS8892554B2Automatic word-cloud generation
Publication Date: 2014.11.18 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US8892554B2 patent drawing
  • US8892554B2 patent drawing
  • US8892554B2 patent drawing

AI summary

Method, system, and computer program product for automatic generation of a word-cloud for a content item are provided. The method includes: extracting terms from a content item using statistical selection criteria; weighting a term by a probability that the term is used as a tag; and generating a visual representation of terms with enhanced representation of terms according to the weighting. Weighting a term by a probability that the term is used as a tag may include determining the relative frequency of the term in a folksonomy of tag terms for a domain.