A scene data labeling method based on a multi-modal large model
By combining multimodal large models with image and text prompts for automatic scene recognition and annotation, the problems of high annotation cost, low efficiency and limited recognition capability in existing technologies are solved, and efficient and accurate scene data annotation and database construction are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NAT AUTOMOBILE UNIV SPACE-TIME TECH (ANQING) CO LTD
- Filing Date
- 2026-03-17
- Publication Date
- 2026-07-10
AI Technical Summary
In existing technologies, data annotation for autonomous driving scenarios relies on manual annotation or a single visual model, which results in high annotation costs, low efficiency, weak model generalization ability, difficulty in meeting the needs of large-scale dataset construction and rapid iteration, and limited recognition ability in complex road environments.
A multimodal large model is used to jointly understand image information and text prompts. By constructing multimodal input data and using the multimodal large model for scene semantic reasoning, a closed-loop optimization mechanism is formed by combining manual quality inspection and model fine-tuning to achieve automatic scene recognition and annotation.
It improves the efficiency of scene data annotation, reduces labor costs, enhances the semantic understanding of complex road environments, has good scene scalability and recognition accuracy, and builds a high-quality scene database.
Smart Images

Figure FT_1 
Figure FT_2