A scene data labeling method based on a multi-modal large model

By combining multimodal large models with image and text prompts for automatic scene recognition and annotation, the problems of high annotation cost, low efficiency and limited recognition capability in existing technologies are solved, and efficient and accurate scene data annotation and database construction are achieved.

CN122369003APending Publication Date: 2026-07-10NAT AUTOMOBILE UNIV SPACE-TIME TECH (ANQING) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NAT AUTOMOBILE UNIV SPACE-TIME TECH (ANQING) CO LTD
Filing Date
2026-03-17
Publication Date
2026-07-10

AI Technical Summary

Technical Problem

In existing technologies, data annotation for autonomous driving scenarios relies on manual annotation or a single visual model, which results in high annotation costs, low efficiency, weak model generalization ability, difficulty in meeting the needs of large-scale dataset construction and rapid iteration, and limited recognition ability in complex road environments.

Method used

A multimodal large model is used to jointly understand image information and text prompts. By constructing multimodal input data and using the multimodal large model for scene semantic reasoning, a closed-loop optimization mechanism is formed by combining manual quality inspection and model fine-tuning to achieve automatic scene recognition and annotation.

Benefits of technology

It improves the efficiency of scene data annotation, reduces labor costs, enhances the semantic understanding of complex road environments, has good scene scalability and recognition accuracy, and builds a high-quality scene database.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
Patent Text Reader

Abstract

This invention discloses a scene data annotation method based on a multimodal large model, belonging to the field of autonomous driving perception technology. This method acquires continuous road scene image frames collected during vehicle operation, constructs text prompts containing scene category information, and combines the image data and text prompts before inputting them into a multimodal large model for inference processing, thereby achieving automatic recognition and annotation of road scenes. Specifically, in the implementation process, continuous image frame sequences are obtained by extracting frames from video data, and text prompts are constructed using prompt templates to guide the multimodal large model to perform cross-modal semantic understanding of the image content, outputting corresponding scene category labels and confidence information. Subsequently, the model output results are parsed to generate structured annotation data, and the annotation results are reviewed and corrected through manual quality inspection. Simultaneously, samples with recognition errors or low confidence are collected and used for model fine-tuning training to improve the model's ability to recognize complex road environments and long-tail scenes. Combined with examples, the results are clearly presented. This invention achieves automatic recognition and annotation of various road scenes, improves scene data annotation efficiency, reduces manual annotation costs, and enhances the system's semantic understanding ability of complex road environments.
Need to check novelty before this filing date? Find Prior Art