Text-Image Semantic Alignment Model for Automated Classification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data mining solutions for vehicle object detection in ADAS and AD systems face challenges such as insufficient accuracy, generalization capabilities, and inefficiency due to manual data search and labor-intensive sample set establishment, leading to inconsistent and time-consuming model training with potential false marks and missed samples.
Innovation Solution
A method and apparatus utilizing a text-image semantic alignment model to automatically generate an image sample set for a classification model based on a pre-trained model through linear probing, enabling quick adaptation to specific scenarios with improved flexibility and accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual data search and sample set establishment are used, then model training can be performed, but the process is time-consuming and labor-intensive with insufficient accuracy
Solution Approach 1:
The system automatically performs data mining, sample set establishment, and model training without requiring manual intervention. The text-image semantic alignment model autonomously selects images based on target text semantics, automatically labels them, and trains the classification model, eliminating the time-consuming manual processes while maintaining high accuracy.
Solution Approach 2:
The patent replaces manual mechanical data collection and labeling processes with automated computer vision and machine learning systems. The text-image semantic alignment model uses automated semantic understanding and image retrieval mechanisms to substitute for manual search and selection, significantly reducing time consumption while improving consistency and accuracy.
2Reliability
If manual sample set establishment is used, then models can be trained, but the process is inconsistent and prone to false marks and missed samples
Solution Approach 1:
The system incorporates feedback mechanisms where the text-image semantic alignment model continuously refines image selection based on target text semantics. The automated labeling process provides consistent feedback to ensure accurate category assignment, eliminating the inconsistency and false marks associated with manual labeling while maintaining manageable process complexity through structured automated workflows.
3Productivity
If pre-trained models are used with linear probing, then model generation speed increases, but adaptability to specific scenarios may be limited
Solution Approach 1:
The system uses pre-trained text-image semantic alignment models as a foundation, performing preliminary learning of general visual and linguistic patterns. This preliminary action enables rapid adaptation to specific driving scenarios through linear probing, where the pre-trained model is quickly fine-tuned on domain-specific data without requiring training from scratch, thus achieving both speed and adaptability.
Solution Approach 2:
The patent applies linear probing which involves changing model parameters (such as adding task-specific layers or adjusting existing weights) rather than retraining the entire model. This parameter change approach allows the pre-trained model to adapt to specific driving scenarios efficiently, maintaining high generation speed while achieving necessary scenario-specific adaptability through targeted parameter adjustments.
Data Source
AI summary
A method and an apparatus, a device, a vehicle, and a medium for generating a classification model are disclosed. The method for generating a classification model includes (i) acquiring, by a text-image semantic alignment model, a plurality of images associated with a target text, which is indicative of a target scene, (ii) generating an image sample set comprising an image sample labeled with a category by determining which category each of the plurality of images belongs to, and (iii) generating a classification model for the target scene based on the image sample set, wherein the generation of the classification model is based on a pre-trained model utilizing linear probing. In this way, customized classification models can be quickly generated for specific scenes to suit actual needs, providing improved flexibility and accuracy while ensuring the efficiency and stability of the model generation process.


