Image Caption Generation With Sentiment-Aware Keyword Prioritization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing automated caption suggestion systems provide raw and descriptive captions that lack sentimental alignment with user requirements, requiring significant user input and literary skills, and are limited by fixed databases and manual search methods.

Innovation Solution

A method and system that determine impacting categories and sentiments of contextual keywords, group them, reorder for desired impact, generate captions, and prioritize based on user profiles and linguistic values, eliminating the need for manual input and fixed databases.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If automated caption suggestion systems use fixed databases and manual search methods, then they can provide caption suggestions, but the suggestions are raw and lack sentimental alignment with user requirements

Engineering Contradiction:
Improvesentimental alignmentVSAvoiduser effort
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The system automatically analyzes image context, determines impacting categories, groups contextual keywords, and generates prioritized captions without requiring user input or manual search. The caption generation engine self-services by creating sentimentally aligned captions directly from image analysis, eliminating the need for users to interact with fixed databases.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces manual search methods and fixed database queries with an automated machine learning system that analyzes image context and generates captions algorithmically. The system substitutes mechanical database searching with intelligent image analysis and automated caption generation, achieving sentimental alignment without user effort.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Adaptability or versatility

If caption generation requires literary skills and significant user input, then captions can be customized, but it requires significant time and user effort

Engineering Contradiction:
Improvecaption customizationVSAvoidtime taken
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The system automatically generates multiple customized captions with different sentimental impacts without requiring user input or literary skills. The caption generation engine self-services by creating adapted captions tailored to user profiles and image context, eliminating the time users would otherwise spend crafting custom captions.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system performs preliminary analysis of image context, determines impacting categories, and pre-generates multiple prioritized captions before user interaction. By preparing customized caption options in advance based on image analysis and user profiles, the system eliminates the need for users to spend time creating captions from scratch.

Inventive Principle:
Principle #10Preliminary action

3Ease of manufacture

If systems use fixed databases of captions, then search and retrieval are simplified, but the system lacks adaptability to different contexts and user profiles

Engineering Contradiction:
Improvesystem implementationVSAvoidcontext adaptability
Core Design Contradiction:
Ease of manufactureVSAdaptability or versatility

Solution Approach 1:

The system dynamically adapts caption generation to different image contexts and user profiles rather than relying on static fixed databases. The impacting category determination engine and caption generation engine adjust their output based on real-time image analysis and user profile data, making the system versatile while maintaining ease of implementation through automated processes.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system changes parameters such as impacting categories, keyword groups, and caption priorities based on image context and user profiles. By dynamically adjusting these parameters rather than using fixed database entries, the system achieves context adaptability while maintaining implementation simplicity through algorithmic generation.

Inventive Principle:
Principle #35Parameter changes

4Productivity

If automated suggestions are generated quickly, then user time is reduced, but the suggestions may be too raw and lack sentimental quality

Engineering Contradiction:
Improvecaption generation speedVSAvoidsentimental quality
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The system performs preliminary analysis of image context and pre-determines impacting categories and keyword groups before generating captions. This preliminary action enables quick generation of sentimentally aligned captions by having the analytical framework ready in advance, achieving both speed and quality simultaneously.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent replaces manual caption crafting with an automated machine learning system that rapidly generates sentimentally aligned captions through image analysis and algorithmic generation. The system substitutes the time-consuming manual process with intelligent automation that maintains sentimental quality while achieving high productivity.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS12525043B2Method and a system for suggesting at least one caption for an image
Publication Date: 2026.01.13 SAMSUNG ELECTRONICS CO LTD
  • US12525043B2 patent drawing
  • US12525043B2 patent drawing
  • US12525043B2 patent drawing

AI summary

Provided is a method for suggesting a caption for an image, comprising: receiving the image; determining a plurality of impacting categories associated with a plurality of contextual keywords for the image, each of the plurality of impacting categories representing a sentiment associated with the plurality of contextual keywords; grouping the plurality of contextual keywords into a plurality of groups based on the plurality of impacting categories; determining an order associated with the plurality of contextual keywords, based on a pre-determined impacting function; generating at least one caption by processing each contextual keyword, based on the order associated with the plurality of contextual keywords; determining a priority value associated with each of the at least one caption based on information associated with the corresponding caption, a user profile, and the image; and suggesting the caption from the at least one caption based on the priority value.