Fine-Tuned Image-Text Filtering for Noisy Datasets
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing image-text datasets are noisy and of low quality, impacting the performance of foundation models in tasks like image classification and text-to-image retrieval, due to the mismatch between image content and text quality.
Innovation Solution
A teacher Multimodal Language Model (MLM) is used to generate instruction data for fine-tuning a machine learning model on quality scoring tasks, including Image-Text Matching, Object Detail Fulfillment, and Caption Text Quality, to filter high-quality image-text pairs based on ITM, ODF, and CTQ metrics.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If existing image-text datasets are used, then the datasets are large in quantity, but the data quality is low and noisy
Solution Approach 1:
The patent extracts high-quality image-text pairs from large datasets by using a teacher MLM to generate quality scores and filter out noisy or low-quality data, retaining only the most relevant and accurate pairs for training
Solution Approach 2:
The patent changes the parameter of data quality by introducing multiple evaluation metrics (ITM, ODF, CTQ) and using fine-tuned models to reassess and rerank data quality, transforming the dataset from low-quality to high-quality through parameter optimization
2Ease of manufacture
If foundation models are trained on noisy data, then the training process is simple, but the model performance deteriorates in tasks like image classification and retrieval
Solution Approach 1:
The patent performs preliminary data filtering and quality assessment before training by using a teacher MLM to pre-evaluate image-text pairs, ensuring that only high-quality data is used for subsequent model training, thus preventing performance deterioration
Solution Approach 2:
The patent introduces a fine-tuned MLM as an intermediary between the raw noisy data and the final model training process, which acts as a mediator to clean, validate, and prepare the data, thereby improving model performance without complicating the overall training workflow
Data Source
AI summary
The present disclosure describes techniques for filtering image-text data. Instruction data is constructed on a plurality of image-text pair quality scoring tasks. A machine learning model is fine-tuned to an image-text data filter using the constructed instruction data. A quality of each image-text pair from a dataset is evaluated by the fine-tuned machine learning model using a plurality of metrics. The plurality of metrics comprises an Image-Text Matching (ITM) metric, an Object Detail Fulfillment (ODF) metric, and a Caption Text Quality (CTQ) metric. High-quality image-text pairs are selected from the dataset based on one or more of the plurality of metrics.


