LLM-Based Image Captioning for Content Recommendation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing systems struggle to effectively recommend non-text content items, such as images, from large corpora to users, due to the difficulty in determining relevant content items from billions of options.
Innovation Solution
The system generates content item captions for selected and provided content items, processes these captions using a Large Language Model (LLM) to produce a narrative description, and then uses this description as a text-based request to identify and return recommended content items.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If traditional search or recommendation systems are used to find relevant content items from billions of options, then the system can maintain a large corpus of content items, but the difficulty of determining relevant content items increases significantly
Solution Approach 1:
The patent introduces image captions as an intermediary representation between images and text queries. These captions serve as a bridge that enables semantic matching between visual content and user queries, making the determination of relevant content items from billions of options feasible through text-based comparison rather than direct image analysis
Solution Approach 2:
The patent replaces traditional image-based or feature-based matching mechanisms with text-based semantic search. By converting images to text captions and using language models, the system substitutes complex image processing and comparison operations with more efficient text-based semantic analysis, significantly reducing the difficulty of determining relevance
2Ease of operation
If text-based queries are used to search for non-text content items, then the system can provide focused search results, but the system struggles to effectively recommend non-text content items from large corpora
Solution Approach 1:
The patent transforms the representation parameters of non-text content items by generating text captions that describe visual content. This parameter change from visual features to text semantics enables the use of text-based queries to effectively search and recommend non-text content, maintaining both query simplicity and recommendation effectiveness
Solution Approach 2:
The patent makes the text caption generation system universal by applying it to various types of non-text content items (images, videos, etc.). This allows the same text-based query interface to work across different content types, providing both focused search results and effective recommendations through a unified system
Data Source
AI summary
Disclosed are systems and methods that process a selection of a plurality of content items through a Large Language Model (“LLM”) and determine, based on that processing, other content items to present or recommend. For example, one or more non-text content item may be processed to generate one or more captions descriptive of the non-text content item(s). The captions may then be processed by a LLM to determine a narrative description of the content items and the narrative description may be used as a text-based query to determine recommended content items.


