LLM/VLM Dataset Generation for Automated Multi-Modal Labeling
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The challenge of data scarcity and the complexity of labeling diverse, high-quality data for multi-modal machine learning models, particularly in domain-specific applications, leads to suboptimal performance and reliability, especially when integrating images and text.
Innovation Solution
A dataset generation system using large language models (LLM) and vision language models (VLM) to iteratively generate and validate datasets, comprising initial queries, reference items, and target items, with cyclical processes to improve model performance.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If domain expertise is used to label data, then label quality is improved, but resource consumption and time requirements increase
Solution Approach 1:
The system uses pre-trained LLMs and VLMs as templates to generate labels automatically, copying their knowledge and reasoning capabilities to produce high-quality labels without manual domain expertise intervention
Solution Approach 2:
The dataset generation system performs self-labeling by automatically generating queries, selecting items, and creating labels through iterative processes without requiring external domain experts for each labeling task
2Adaptability or versatility
If diverse data types are integrated for multi-modal models, then model versatility is improved, but labeling complexity increases
Solution Approach 1:
The system employs multi-functional LLMs and VLMs that can handle multiple data types (text, images) and perform multiple tasks (query generation, item selection, label creation) through a unified framework, reducing the need for separate specialized labeling processes for each modality
Solution Approach 2:
The patent introduces an intermediary automated labeling system that mediates between raw diverse data and model training requirements, translating various data types into standardized labels through LLM/VLM-based processing
3Reliability
If more training data is acquired, then model performance is improved, but data acquisition costs increase
Solution Approach 1:
Instead of acquiring large volumes of real-world labeled data through expensive manual processes, the system copies the labeling capabilities of domain experts into automated LLM/VLM systems that can generate unlimited synthetic labeled data
Solution Approach 2:
The system performs preliminary automated labeling and validation before model training, preparing extensive training datasets in advance through iterative query generation and item selection processes
Data Source
AI summary
In general, a model development platform that includes a dataset generation system is disclosed. The dataset generation system may generate a dataset, which may be useable to train or validate performance of a multi-modal machine learning model. To generate a sample of the dataset, the dataset generation system may use an initial query to search for a reference item and target item. The dataset generation system may use one or more of a large language model or vision language model to generate a follow-up query. In some embodiments, the multi-modal machine learning model may validate the dataset, which may then be used to train the multi-modal machine learning model, which may then be used to generate a subsequent dataset creating, in some embodiments, a cyclical process for improving the machine learning model.


