Text-Image Semantic Alignment Model for Automated Classification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current data mining solutions for vehicle object detection in ADAS and AD systems face challenges such as insufficient accuracy, generalization capabilities, and inefficiency due to manual data search and labor-intensive sample set establishment, leading to inconsistent and time-consuming model training with potential false marks and missed samples.

Innovation Solution

A method and apparatus utilizing a text-image semantic alignment model to automatically generate an image sample set for a classification model based on a pre-trained model through linear probing, enabling quick adaptation to specific scenarios with improved flexibility and accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual data search and sample set establishment are used, then model training can be performed, but the process is time-consuming and labor-intensive with insufficient accuracy

Engineering Contradiction:
Improvemodel accuracyVSAvoidtraining time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system automatically performs data mining, sample set establishment, and model training without requiring manual intervention. The text-image semantic alignment model autonomously selects images based on target text semantics, automatically labels them, and trains the classification model, eliminating the time-consuming manual processes while maintaining high accuracy.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces manual mechanical data collection and labeling processes with automated computer vision and machine learning systems. The text-image semantic alignment model uses automated semantic understanding and image retrieval mechanisms to substitute for manual search and selection, significantly reducing time consumption while improving consistency and accuracy.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Reliability

If manual sample set establishment is used, then models can be trained, but the process is inconsistent and prone to false marks and missed samples

Engineering Contradiction:
Improvetraining consistencyVSAvoidprocess complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system incorporates feedback mechanisms where the text-image semantic alignment model continuously refines image selection based on target text semantics. The automated labeling process provides consistent feedback to ensure accurate category assignment, eliminating the inconsistency and false marks associated with manual labeling while maintaining manageable process complexity through structured automated workflows.

Inventive Principle:
Principle #23Feedback

3Productivity

If pre-trained models are used with linear probing, then model generation speed increases, but adaptability to specific scenarios may be limited

Engineering Contradiction:
Improvemodel generation speedVSAvoidscenario adaptability
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The system uses pre-trained text-image semantic alignment models as a foundation, performing preliminary learning of general visual and linguistic patterns. This preliminary action enables rapid adaptation to specific driving scenarios through linear probing, where the pre-trained model is quickly fine-tuned on domain-specific data without requiring training from scratch, thus achieving both speed and adaptability.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent applies linear probing which involves changing model parameters (such as adding task-specific layers or adjusting existing weights) rather than retraining the entire model. This parameter change approach allows the pre-trained model to adapt to specific driving scenarios efficiently, maintaining high generation speed while achieving necessary scenario-specific adaptability through targeted parameter adjustments.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20250209793A1Method and Apparatus, Device, Vehicle, and Medium for Generating Classification Model
Publication Date: 2025.06.26 ROBERT BOSCH GMBH
  • US20250209793A1 patent drawing
  • US20250209793A1 patent drawing
  • US20250209793A1 patent drawing

AI summary

A method and an apparatus, a device, a vehicle, and a medium for generating a classification model are disclosed. The method for generating a classification model includes (i) acquiring, by a text-image semantic alignment model, a plurality of images associated with a target text, which is indicative of a target scene, (ii) generating an image sample set comprising an image sample labeled with a category by determining which category each of the plurality of images belongs to, and (iii) generating a classification model for the target scene based on the image sample set, wherein the generation of the classification model is based on a pre-trained model utilizing linear probing. In this way, customized classification models can be quickly generated for specific scenes to suit actual needs, providing improved flexibility and accuracy while ensuring the efficiency and stability of the model generation process.