Prompt-Guided Multi-Modal Recognition Without Task-Specific Training
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current pre-trained multi-modal large models require data labeling and model training to perform specific multi-modal recognition tasks, which are time-consuming, labor-intensive, and costly in terms of computing power.
Innovation Solution
A data processing method that utilizes pre-trained multi-modal models and prompts to generate text descriptions of image data, allowing direct input into a large language model for recognition without data labeling or model training, leveraging the pre-trained models' capabilities across various service scenarios.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If data labeling and model training are performed to implement specific multi-modal recognition tasks, then recognition accuracy is improved, but computing power costs and time costs increase significantly
Solution Approach 1:
The patent applies preliminary action by using pre-trained multi-modal large models that have already been trained on extensive data beforehand. These pre-trained models can directly perform specific multi-modal recognition tasks without requiring additional data labeling and model training, thus achieving high recognition accuracy while significantly reducing time costs and computational resource consumption.
2Measurement precision
If data labeling and model training are performed to implement specific multi-modal recognition tasks, then recognition accuracy is improved, but computing power costs increase significantly
Solution Approach 1:
The patent applies preliminary action by using pre-trained multi-modal large models that have already been trained on extensive data beforehand. These pre-trained models can directly perform specific multi-modal recognition tasks without requiring additional data labeling and model training, thus achieving high recognition accuracy while significantly reducing time costs and computational resource consumption.
3Measurement precision
If data labeling is performed for specific multi-modal recognition tasks, then task-specific accuracy is improved, but labor intensity increases
Solution Approach 1:
The patent applies self-service by using pre-trained multi-modal large models that possess general multi-modal understanding capabilities. These models can directly handle specific recognition tasks without requiring human experts to manually label data, thereby maintaining task-specific accuracy while dramatically reducing labor intensity and making the system easier to operate.
4Productivity
If pre-trained multi-modal models are used without data labeling, then execution efficiency is improved, but direct implementation of specific tasks becomes possible
Solution Approach 1:
The patent applies universality by designing a unified framework where pre-trained multi-modal large models serve multiple purposes. The same pre-trained models can directly implement various specific multi-modal recognition tasks (such as classification, visual question answering, etc.) without requiring task-specific data labeling or fine-tuning, thus achieving both high execution efficiency and broad task adaptability through a single multi-functional system.
Data Source
AI summary
In a data processing method, input data is acquired. The input data includes image data. A label of the image data is acquired through a first multi-modal model and a word list. The label identifies at least one element present in the image data. The word list is defined for a recognition task and includes N words. N is a positive integer. Through a second multi-modal model and a second prompt, a text description of the image data is acquired. The second prompt controls generation of an image content description corresponding to the recognition task. First text information from the label and the text description is generated based on a target prompt. The target prompt is defined for the recognition task. The first text information is input into a large language model. A recognition result of the input data is output.


