Image Captioning Dataset Generation for Low-Resource Languages
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for constructing image captioning data require significant human intervention and resources, making it costly and difficult to obtain captions for languages with few users.
Innovation Solution
An apparatus and method that utilizes processors and memory to paraphrase and translate input sentences into multiple expressions and languages, generating a captioning dataset with minimal human intervention.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If human annotators directly input captions into computing devices, then captioning data can be obtained, but it requires a lot of cost and time
Solution Approach 1:
The system performs preliminary actions by using AI models to pre-generate captioning data for multiple languages before human annotators review it. This reduces the time required for data construction by handling the initial data generation automatically, while human annotators only need to perform quality verification and correction on the pre-generated data.
2Reliability
If human annotators directly input captions, then captioning data can be obtained, but it requires significant human resources and costs
Solution Approach 1:
The system introduces AI models as intermediaries between the image input and final captioning data. These models automatically generate multilingual captions, reducing the need for human annotators to directly create captions from scratch. Human annotators serve as a secondary intermediary layer for quality control, creating a balanced automation-human collaboration system.
3Adaptability or versatility
If human annotators create captions for each language separately, then language-specific captioning data can be obtained, but it is not easy to obtain captioning data for languages with few users
Solution Approach 1:
The system implements universality by using a single AI model that can generate captioning data for multiple languages simultaneously. Instead of requiring separate human annotation teams for each language, the model performs the function of creating captions across many languages with a single automated process, making it easy to obtain data for languages with few users.
Data Source
AI summary
An apparatus for constructing captioning data according to an embodiment is provided with one or more processors and a memory storing one or more programs executed by the one or more processors, and includes an input module configured to acquire input sentences for an image for which captions are to be acquired, and a generation module configured to generate a captioning dataset by paraphrasing and translating the input sentences to generate a plurality of paraphrased and translated sentences and using the plurality of generated paraphrased and translated sentences as a set.


