LLM-Based Image Captioning for Content Recommendation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing systems struggle to effectively recommend non-text content items, such as images, from large corpora to users, due to the difficulty in determining relevant content items from billions of options.

Innovation Solution

The system generates content item captions for selected and provided content items, processes these captions using a Large Language Model (LLM) to produce a narrative description, and then uses this description as a text-based request to identify and return recommended content items.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If traditional search or recommendation systems are used to find relevant content items from billions of options, then the system can maintain a large corpus of content items, but the difficulty of determining relevant content items increases significantly

Engineering Contradiction:
Improvecorpus sizeVSAvoidrelevance determination difficulty
Core Design Contradiction:
Quantity of substanceVSDifficulty of detecting and measuring

Solution Approach 1:

The patent introduces image captions as an intermediary representation between images and text queries. These captions serve as a bridge that enables semantic matching between visual content and user queries, making the determination of relevant content items from billions of options feasible through text-based comparison rather than direct image analysis

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent replaces traditional image-based or feature-based matching mechanisms with text-based semantic search. By converting images to text captions and using language models, the system substitutes complex image processing and comparison operations with more efficient text-based semantic analysis, significantly reducing the difficulty of determining relevance

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Ease of operation

If text-based queries are used to search for non-text content items, then the system can provide focused search results, but the system struggles to effectively recommend non-text content items from large corpora

Engineering Contradiction:
Improvequery simplicityVSAvoidrecommendation effectiveness
Core Design Contradiction:
Ease of operationVSProductivity

Solution Approach 1:

The patent transforms the representation parameters of non-text content items by generating text captions that describe visual content. This parameter change from visual features to text semantics enables the use of text-based queries to effectively search and recommend non-text content, maintaining both query simplicity and recommendation effectiveness

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent makes the text caption generation system universal by applying it to various types of non-text content items (images, videos, etc.). This allows the same text-based query interface to work across different content types, providing both focused search results and effective recommendations through a unified system

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20250139353A1Identifying image based content items using a large language model
  • US20250139353A1 patent drawing
  • US20250139353A1 patent drawing
  • US20250139353A1 patent drawing

AI summary

Disclosed are systems and methods that process a selection of a plurality of content items through a Large Language Model (“LLM”) and determine, based on that processing, other content items to present or recommend. For example, one or more non-text content item may be processed to generate one or more captions descriptive of the non-text content item(s). The captions may then be processed by a LLM to determine a narrative description of the content items and the narrative description may be used as a text-based query to determine recommended content items.