Image Captioning Dataset Generation for Low-Resource Languages

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for constructing image captioning data require significant human intervention and resources, making it costly and difficult to obtain captions for languages with few users.

Innovation Solution

An apparatus and method that utilizes processors and memory to paraphrase and translate input sentences into multiple expressions and languages, generating a captioning dataset with minimal human intervention.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If human annotators directly input captions into computing devices, then captioning data can be obtained, but it requires a lot of cost and time

Engineering Contradiction:
Improvecaptioning data qualityVSAvoiddata construction time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system performs preliminary actions by using AI models to pre-generate captioning data for multiple languages before human annotators review it. This reduces the time required for data construction by handling the initial data generation automatically, while human annotators only need to perform quality verification and correction on the pre-generated data.

Inventive Principle:
Principle #10Preliminary action

2Reliability

If human annotators directly input captions, then captioning data can be obtained, but it requires significant human resources and costs

Engineering Contradiction:
Improvecaptioning data qualityVSAvoidhuman intervention level
Core Design Contradiction:
ReliabilityVSExtent of automation

Solution Approach 1:

The system introduces AI models as intermediaries between the image input and final captioning data. These models automatically generate multilingual captions, reducing the need for human annotators to directly create captions from scratch. Human annotators serve as a secondary intermediary layer for quality control, creating a balanced automation-human collaboration system.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Adaptability or versatility

If human annotators create captions for each language separately, then language-specific captioning data can be obtained, but it is not easy to obtain captioning data for languages with few users

Engineering Contradiction:
Improvelanguage coverageVSAvoiddata construction ease
Core Design Contradiction:
Adaptability or versatilityVSEase of manufacture

Solution Approach 1:

The system implements universality by using a single AI model that can generate captioning data for multiple languages simultaneously. Instead of requiring separate human annotation teams for each language, the model performs the function of creating captions across many languages with a single automated process, making it easy to obtain data for languages with few users.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20260044553A1Apparatus and method for constructing captioning data for images
Publication Date: 2026.02.12 CHUNG ANG UNIV IND ACADEMIC COOP FOUND
  • US20260044553A1 patent drawing
  • US20260044553A1 patent drawing
  • US20260044553A1 patent drawing

AI summary

An apparatus for constructing captioning data according to an embodiment is provided with one or more processors and a memory storing one or more programs executed by the one or more processors, and includes an input module configured to acquire input sentences for an image for which captions are to be acquired, and a generation module configured to generate a captioning dataset by paraphrasing and translating the input sentences to generate a plurality of paraphrased and translated sentences and using the plurality of generated paraphrased and translated sentences as a set.