VQA Model Training via Non-VQA Data Conversion
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The high cost and complexity of preparing diverse training datasets for visual question answering (VQA) models limit their ability to achieve high accuracy, as they require numerous variations of images, questions, and answers, making efficient and cost-effective learning challenging.
Innovation Solution
A machine learning apparatus that converts non-VQA format samples into VQA format samples by generating question and answer texts based on object labels, allowing the training of VQA models using a combination of VQA and non-VQA samples, thereby increasing the variety of training data and improving model accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a training data set with wide variations of images, questions and answers is prepared, then the accuracy of the statistical model is improved, but the cost and complexity of data preparation increases enormously
Solution Approach 1:
The patent generates synthetic VQA samples by copying and transforming existing non-VQA data (image-label pairs) into VQA format (image-question-answer triples). This allows the model to be trained on diverse VQA-like data without requiring manual creation of each sample, thereby reducing data preparation complexity while maintaining training data diversity for high accuracy
Solution Approach 2:
The system uses existing non-VQA data (such as image classification data) to automatically generate VQA training samples through a conversion process. This self-service approach leverages already-available data resources to create the needed training material, eliminating the need for expensive manual annotation and question generation while still producing diverse training samples
2Device complexity
If a statistical model is trained with a training data set with less variations to reduce cost, then the cost of data preparation is reduced, but the accuracy of the statistical model cannot be generated
Solution Approach 1:
The patent transforms non-VQA data into VQA format by changing the parameter structure of the data. Non-VQA data consisting of image-label pairs is converted into VQA format with generated questions and answers, thereby altering the data structure to match the training requirements while using the same underlying image data. This allows diverse training samples to be created from existing data without additional cost
3Adaptability or versatility
If hundreds of thousands of questions with respect to several tens of thousands of images are prepared, then the statistical model can support specific animals, plants or vehicles, but the cost of preparing the training data set increases enormously
Solution Approach 1:
The patent creates synthetic VQA samples by copying existing non-VQA data and transforming it into the required VQA format. This allows the model to learn about specific objects (animals, plants, vehicles) from existing image-label data without requiring manual creation of targeted training samples, thereby reducing the quantity of data needed while maintaining adaptability to specific object categories
Data Source
AI summary
According to one embodiment, a machine learning apparatus includes a processing circuit. The processing circuit generates a training sample in a VQA format regarding a VQA task based on a sample in a non-VQA format. The training sample in the VQA format includes a combination of an object, a question text regarding the object and an answer text in response to the question text as elements, and the sample in the non-VQA format includes a combination of an object and a label related to the object as elements. The processing circuit trains a statistical model of the VQA task based on the generated training sample in the VQA format.


