Non-textual Hashtag Generation via Visual Embeddings

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing hashtag systems are limited in their ability to describe non-textual content effectively, as they rely on textual information and struggle to capture the essence of content types like images, videos, and audio, which cannot be adequately represented using text-based keywords.

Innovation Solution

A computer-implemented tagging system that uses a joint embedding model to perform contextual analysis of non-textual content, allowing for the automatic generation and recommendation of non-textual hashtags that include both textual and non-textual components, enabling better description and search functionality across various media types.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If textual hashtags are used to tag non-textual content, then the system maintains simplicity and ease of operation, but the descriptiveness and accuracy of content representation deteriorates

Engineering Contradiction:
Improveease of hashtag creationVSAvoiddescriptiveness of content
Core Design Contradiction:
Ease of operationVSLoss of information

Solution Approach 1:

The patent introduces visual embeddings as an intermediary representation layer between non-textual content and textual hashtags. The system converts non-textual content (images, videos, audio) into visual embeddings that capture essential visual features, then uses these embeddings to generate or select appropriate textual hashtags. This intermediary process preserves the simplicity of textual hashtag interfaces while significantly improving content representation accuracy by leveraging learned visual features rather than direct text description.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If non-textual hashtags are automatically generated using contextual analysis, then the descriptiveness and accuracy of content tagging improves, but the system complexity and computational requirements increase

Engineering Contradiction:
Improveaccuracy of content descriptionVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent implements preliminary action by pre-training visual embedding models on large datasets of non-textual content before deployment. The contextual analysis framework and attribute identification capabilities are established in advance through model training, allowing the system to quickly generate accurate hashtags during actual use without requiring complex real-time analysis infrastructure. This pre-computation approach reduces operational complexity while maintaining high accuracy.

Inventive Principle:
Principle #10Preliminary action

3Loss of information

If visual embeddings are used to represent non-textual content, then the ability to capture content essence improves, but the computational resources and processing time required increase

Engineering Contradiction:
Improvecapture of content essenceVSAvoidcomputational energy consumption
Core Design Contradiction:
Loss of informationVSUse of energy by moving object

Solution Approach 1:

The patent applies parameter changes by allowing dynamic adjustment of embedding dimensionality and complexity based on specific application requirements. The system can configure visual embeddings with different parameter settings (e.g., embedding size, feature extraction depth) to balance between capturing content essence and computational efficiency. This flexibility enables the same framework to serve both high-accuracy research applications and resource-constrained deployment scenarios.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20230161775A1Non-textual hashtag creation for non-textual content
Publication Date: 2023.05.25 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US20230161775A1 patent drawing
  • US20230161775A1 patent drawing
  • US20230161775A1 patent drawing

AI summary

A computer-implemented process within a tagging system configured to be executed on a computer hardware system includes the following operations. An identification of non-textual content being accessed by the user is received from a client tagging module within a client device associated with a user. A contextual analysis of the non-textual content is performed, using an object identification engine of the tagging system, to identify attributes of the non-textual content. An identification of the non-textual hashtag is stored as a data structure in association with the non-textual content and at least one of the attributes and the non-textual content. A search is performed for additional non-textual content related to the non-textual content based upon a selection of the non-textual hashtag.