Fine-Tuned Image-Text Filtering for Noisy Datasets

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing image-text datasets are noisy and of low quality, impacting the performance of foundation models in tasks like image classification and text-to-image retrieval, due to the mismatch between image content and text quality.

Innovation Solution

A teacher Multimodal Language Model (MLM) is used to generate instruction data for fine-tuning a machine learning model on quality scoring tasks, including Image-Text Matching, Object Detail Fulfillment, and Caption Text Quality, to filter high-quality image-text pairs based on ITM, ODF, and CTQ metrics.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If existing image-text datasets are used, then the datasets are large in quantity, but the data quality is low and noisy

Engineering Contradiction:
Improvedataset sizeVSAvoiddata quality
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

The patent extracts high-quality image-text pairs from large datasets by using a teacher MLM to generate quality scores and filter out noisy or low-quality data, retaining only the most relevant and accurate pairs for training

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent changes the parameter of data quality by introducing multiple evaluation metrics (ITM, ODF, CTQ) and using fine-tuned models to reassess and rerank data quality, transforming the dataset from low-quality to high-quality through parameter optimization

Inventive Principle:
Principle #35Parameter changes

2Ease of manufacture

If foundation models are trained on noisy data, then the training process is simple, but the model performance deteriorates in tasks like image classification and retrieval

Engineering Contradiction:
Improvetraining simplicityVSAvoidmodel performance
Core Design Contradiction:
Ease of manufactureVSReliability

Solution Approach 1:

The patent performs preliminary data filtering and quality assessment before training by using a teacher MLM to pre-evaluate image-text pairs, ensuring that only high-quality data is used for subsequent model training, thus preventing performance deterioration

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces a fine-tuned MLM as an intermediary between the raw noisy data and the final model training process, which acts as a mediator to clean, validate, and prepare the data, thereby improving model performance without complicating the overall training workflow

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20250278928A1Filtering image-text data using a fine-tuned machine learning model
Publication Date: 2025.09.04 LEMON INC(GB)
  • US20250278928A1 patent drawing
  • US20250278928A1 patent drawing
  • US20250278928A1 patent drawing

AI summary

The present disclosure describes techniques for filtering image-text data. Instruction data is constructed on a plurality of image-text pair quality scoring tasks. A machine learning model is fine-tuned to an image-text data filter using the constructed instruction data. A quality of each image-text pair from a dataset is evaluated by the fine-tuned machine learning model using a plurality of metrics. The plurality of metrics comprises an Image-Text Matching (ITM) metric, an Object Detail Fulfillment (ODF) metric, and a Caption Text Quality (CTQ) metric. High-quality image-text pairs are selected from the dataset based on one or more of the plurality of metrics.