Auxiliary-Guided Image Caption Generation for Product Imagery

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Manually generating image captions is time-consuming and inefficient, making it difficult to meet the demands of large-scale production, especially in e-commerce applications where product images often contain multiple types of information, and simply displaying the image does not allow users to quickly grasp the intended focus.

Innovation Solution

An end-to-end method for generating image captions by obtaining an image and auxiliary caption information, determining image and auxiliary features, and generating target captions that include the name information of the main object, utilizing a cloud-based server with computing nodes to automate the process.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual caption generation is used, then caption quality can be controlled, but time consumption and labor intensity increase significantly

Engineering Contradiction:
Improvecaption qualityVSAvoidtime consumption
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent replaces the manual mechanical process of writing captions with an automated AI-based system. The system uses image processing algorithms to extract object information, category, and attributes from product images, then generates captions automatically without human intervention, thereby eliminating time consumption while maintaining quality through structured data extraction and generation processes

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables self-service by automatically processing images and generating captions independently without requiring manual input. The automated system extracts information from images, determines object characteristics, and generates appropriate captions on its own, freeing users from manual caption writing tasks while maintaining consistent quality through programmed algorithms

Inventive Principle:
Principle #25Self-service

2Measurement precision

If manual caption generation is used, then accuracy can be ensured, but productivity decreases

Engineering Contradiction:
Improvecaption accuracyVSAvoidproduction efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent substitutes manual caption writing with an automated AI system that processes images and generates captions efficiently. The system uses computer vision algorithms to accurately identify objects, categories, and attributes from product images, then generates captions automatically, thereby achieving both high accuracy and high productivity without the trade-off present in manual processes

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables continuous automated caption generation without interruptions. Unlike manual processes that require human attention and have variable speeds, the automated system continuously processes images and generates captions at a constant high speed, maintaining both accuracy through structured algorithms and productivity through uninterrupted operation

Inventive Principle:
Principle #20Continuity of useful action

3Ease of operation

If only product images are displayed, then simplicity is maintained, but user understanding of image focus is reduced

Engineering Contradiction:
Improvedisplay simplicityVSAvoidimage focus information
Core Design Contradiction:
Ease of operationVSLoss of information

Solution Approach 1:

The patent introduces captions as an intermediary element between the product image and the user. The captions serve as a mediator that bridges the gap by providing textual descriptions of the main object, category, and attributes, thereby enhancing user understanding of image focus without complicating the display interface. Users can quickly grasp key information by reading the caption alongside the image

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20250239094A1Image caption generation method, device, and computer storage medium
Publication Date: 2025.07.24 HANGZHOU ALIBABA INT INTERNET IND CO LTD
  • US20250239094A1 patent drawing
  • US20250239094A1 patent drawing
  • US20250239094A1 patent drawing

AI summary

The present application provides a method for device, and computer storage medium for generating an image caption. The method comprises: obtaining an image to be processed and auxiliary caption information, wherein the image to be processed includes a main object, and the auxiliary caption information includes at least one of the following: name information corresponding to the main object, object category corresponding to the main object, an object attribute corresponding to the main object, and an image tag corresponding to the image to be processed; determining an image feature corresponding to the image to be processed, and an auxiliary feature corresponding to the auxiliary caption information; generating the caption based on the image feature and the auxiliary feature to obtain a target caption corresponding to the image to be processed, wherein the target caption includes the name information of the main object.