LLM/VLM Dataset Generation for Automated Multi-Modal Labeling

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The challenge of data scarcity and the complexity of labeling diverse, high-quality data for multi-modal machine learning models, particularly in domain-specific applications, leads to suboptimal performance and reliability, especially when integrating images and text.

Innovation Solution

A dataset generation system using large language models (LLM) and vision language models (VLM) to iteratively generate and validate datasets, comprising initial queries, reference items, and target items, with cyclical processes to improve model performance.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If domain expertise is used to label data, then label quality is improved, but resource consumption and time requirements increase

Engineering Contradiction:
Improvelabel qualityVSAvoidlabeling time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system uses pre-trained LLMs and VLMs as templates to generate labels automatically, copying their knowledge and reasoning capabilities to produce high-quality labels without manual domain expertise intervention

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The dataset generation system performs self-labeling by automatically generating queries, selecting items, and creating labels through iterative processes without requiring external domain experts for each labeling task

Inventive Principle:
Principle #25Self-service

2Adaptability or versatility

If diverse data types are integrated for multi-modal models, then model versatility is improved, but labeling complexity increases

Engineering Contradiction:
Improvemodel versatilityVSAvoidlabeling complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system employs multi-functional LLMs and VLMs that can handle multiple data types (text, images) and perform multiple tasks (query generation, item selection, label creation) through a unified framework, reducing the need for separate specialized labeling processes for each modality

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent introduces an intermediary automated labeling system that mediates between raw diverse data and model training requirements, translating various data types into standardized labels through LLM/VLM-based processing

Inventive Principle:
Principle #24Intermediary (Mediator)

3Reliability

If more training data is acquired, then model performance is improved, but data acquisition costs increase

Engineering Contradiction:
Improvemodel performanceVSAvoiddata acquisition resources
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

Instead of acquiring large volumes of real-world labeled data through expensive manual processes, the system copies the labeling capabilities of domain experts into automated LLM/VLM systems that can generate unlimited synthetic labeled data

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system performs preliminary automated labeling and validation before model training, preparing extensive training datasets in advance through iterative query generation and item selection processes

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20250335961A1Dataset generation pipeline using large language models and vision language models
Publication Date: 2025.10.30 TARGET BRANDS INC
  • US20250335961A1 patent drawing
  • US20250335961A1 patent drawing
  • US20250335961A1 patent drawing

AI summary

In general, a model development platform that includes a dataset generation system is disclosed. The dataset generation system may generate a dataset, which may be useable to train or validate performance of a multi-modal machine learning model. To generate a sample of the dataset, the dataset generation system may use an initial query to search for a reference item and target item. The dataset generation system may use one or more of a large language model or vision language model to generate a follow-up query. In some embodiments, the multi-modal machine learning model may validate the dataset, which may then be used to train the multi-modal machine learning model, which may then be used to generate a subsequent dataset creating, in some embodiments, a cyclical process for improving the machine learning model.