Man-machine collaborative data labeling and cleaning method based on multi-modal large language model

By applying a human-machine collaboration method of multimodal large language model in multimodal data annotation and cleaning, the accuracy and efficiency problems of multimodal data annotation and cleaning in the prior art are solved, and fast and accurate multi-task data annotation is achieved, reducing labor costs.

CN120086754APending Publication Date: 2025-06-03BEIJING ZHONGKE RUIJIAN TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202411986176.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

The prior art is difficult to realize accurate data annotation and cleaning in multimodal data, especially in the face of multitasking, and the manual checksum review of algorithm models depend on is insufficient, resulting in labeling accuracy problems.

Method used

The human-computer collaborative data annotation and cleaning method based on the multimodal large language model is adopted. By obtaining the text prompt words of the data annotation instruction, the data characteristics of the data to be marked are extracted, and mapped to the text space, the large language model is input to generate data annotation information, and combined with manual review to improve the annotation accuracy.

Benefits of technology

It realizes fast and accurate labeling of multimodal data, reduces labor costs, improves the accuracy and consistency of labeling results, and is suitable for complex multi-task data labeling tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120086754A_ABST
    Figure CN120086754A_ABST
Patent Text Reader

Abstract

The invention relates to a man-machine collaborative data labeling and cleaning method based on a multi-modal large language model, which comprises the following steps of: acquiring a text cue word of a data labeling instruction, and extracting a text feature of the text cue word through a text encoder; a feature extraction module is utilized to extract data features associated with the annotation object in the data annotation instruction in the to-be-annotated data; mapping the data features to a text space to obtain a data feature text; inputting the text feature and the data feature text into a large language model, understanding the to-be-labeled data, and generating data labeling information of the to-be-labeled data corresponding to the text cue word; a data annotation module is used for completing data annotation of the to-be-annotated data, and an algorithm annotation result is obtained; displaying the algorithm labeling result to a worker, and obtaining an accuracy judgment conclusion of the worker on the algorithm labeling result; and based on the accuracy judgment conclusion of the algorithm labeling result, cleaning the algorithm labeling result to obtain a data labeling result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a human - machine collaborative data annotation and cleaning method based on a multi - modal large - language model, which is applicable to the field of data marking technology. Background Art

[0002] There is a large amount of data of various modal types such as audio, video, images, and text on the Internet, and this data is very messy. How to accurately label this data is a big problem. Current technical solutions can be divided into two categories, namely, manually cleaning and annotating data, and training a single - task deep - learning model for data cleaning and annotation. Using the first method of manual cleaning and annotation undoubtedly increases a lot of labor costs. At the same time, since it is manual annotation, the annotation standard depends more on the subjective consciousness of the annotator, and there are also prone to situations such as mislabeling, missing labeling, and inconsistent annotation standards. Using the second method, a large amount of single - modal data such as audio, images, and text manually annotated is used to train a single - task deep - learning algorithm model, and data cleaning and annotation for a specific task are completed based on the trained deep - learning algorithm model. However, this method has great limitations and cannot complete complex annotation tasks.

[0003] The above two methods have the following defects: The first is that in the face of multiple tasks, it is difficult to clean and annotate data. For example, in video annotation tasks, how to understand a video from subtitles, human voices, background audio, multi - language understanding, image characters, image scenes, and image elements, and give accurate labels, as well as clean and annotate multiple tasks such as object detection, image segmentation, image classification, and speech recognition. The second is the problem of annotation accuracy. If only relying on the algorithm model to achieve data cleaning and annotation without manual verification and review, the obtained annotated data set is very likely to be incorrect, and training the algorithm model with mislabeled data will directly reduce the ability of the algorithm model. Summary of the Invention

[0004] The technical problem to be solved by the present invention is: To solve the above - mentioned technical problems, the present invention provides a human - machine collaborative data annotation and cleaning method based on a multi - modal large - language model.

[0005] The technical solution adopted by the present invention is: A human - machine collaborative data annotation and cleaning method based on a multi - modal large - language model, including:

[0006] Obtain the text prompt of the data annotation instruction, and extract the text features of the text prompt through a text encoder;

[0007] Based on the type of the data to be annotated, use the corresponding feature extraction module of this type to extract the data features associated with the annotation object in the data to be annotated of this type;

[0008] Map the data features of the data to be annotated to the text space to obtain the data feature text corresponding to the data features of the data to be annotated;

[0009] Input the text features of the text prompt and the data feature text of the data to be annotated into the large language model, and use the large language model to understand the data to be annotated and generate the data annotation information corresponding to the text prompt for this data to be annotated;

[0010] Based on the data annotation information, use the data annotation module to complete the data annotation of the data to be annotated to obtain the algorithm annotation result;

[0011] Display the algorithm annotation result to the staff and obtain the judgment conclusion on the accuracy of the algorithm annotation result by the staff;

[0012] Based on the judgment conclusion on the accuracy of the algorithm annotation result, clean the algorithm annotation result to obtain the data annotation result.

[0013] Based on the type of the data to be annotated, use the corresponding feature extraction module of this type to extract the data features associated with the annotation object in the data to be annotated of this type, including:

[0014] If the data to be annotated is image data, use the feature extraction module to extract spatial features from the image data;

[0015] If the data to be annotated is video data, use the feature extraction module to extract spatio-temporal features and / or speech features from the video data.

[0016] A human-computer collaborative data annotation and cleaning device based on a multimodal large language model, having:

[0017] A text feature extraction module, configured to obtain the text prompt of the data annotation instruction and extract the text features of the text prompt through a text encoder;

[0018] A data feature extraction module, configured to, based on the type of the data to be annotated, use the corresponding feature extraction module of this type to extract the data features associated with the annotation object in the data to be annotated of this type;

[0019] A data feature text acquisition module, configured to map the data features of the data to be annotated to the text space to obtain the data feature text corresponding to the data features of the data to be annotated;

[0020] A data annotation information generation module, configured to input the text features of the text prompt and the data feature text of the data to be annotated into the large language model, use the large language model to understand the data to be annotated, and generate the data annotation information corresponding to the text prompt for this data to be annotated;

[0021] The algorithm annotation result generation module is used to complete the data annotation of the data to be annotated by using the data annotation module based on the data annotation information, and obtain the algorithm annotation result;

[0022] The accuracy judgment module is used to display the algorithm annotation result to the staff and obtain the accuracy judgment conclusion of the staff on the algorithm annotation result;

[0023] The data annotation result generation module is used to clean the algorithm annotation result based on the accuracy judgment conclusion of the algorithm annotation result, and obtain the data annotation result.

[0024] A storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above-mentioned human-computer collaborative data annotation and cleaning method based on a multimodal large language model are realized.

[0025] A data annotation device based on a multimodal large language model, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the above-mentioned human-computer collaborative data annotation and cleaning method based on a multimodal large language model are realized.

[0026] A data annotation method based on a multimodal large language model, including:

[0027] Obtain the text prompt words of the data annotation instruction, and extract the text features of the text prompt words through a text encoder;

[0028] Based on the type of the data to be annotated, use the corresponding feature extraction module of this type to extract the data features associated with the annotation object in the data to be annotated of this type;

[0029] Map the data features of the data to be annotated to the text space to obtain the data feature text corresponding to the data features of the data to be annotated;

[0030] Input the text features of the text prompt words and the data feature text of the data to be annotated into the large language model, use the large language model to understand the data to be annotated, and generate the data annotation information corresponding to the text prompt words for this data to be annotated;

[0031] Based on the data annotation information, use the data annotation module to complete the data annotation of the data to be annotated.

[0032] The step of extracting the data features associated with the annotation object in the data to be annotated of this type based on the type of the data to be annotated by using the corresponding feature extraction module of this type includes:

[0033] If the data to be annotated is image data, use the feature extraction module to extract spatial features from the image data;

[0034] If the data to be annotated is video data, the spatio-temporal features and / or speech features are extracted from the video data by using the feature extraction module.

[0035] A data annotation device based on a multimodal large language model, comprising:

[0036] A text feature extraction module, configured to obtain the text prompt of the data annotation instruction, and extract the text features of the text prompt through a text encoder;

[0037] A data feature extraction module, configured to, based on the type of the data to be annotated, use the corresponding feature extraction module of this type to extract the data features associated with the annotation object in the data annotation instruction from the data to be annotated of this type;

[0038] A data feature text acquisition module, configured to map the data features of the data to be annotated into the text space to obtain the data feature text corresponding to the data features of the data to be annotated;

[0039] A data annotation information generation module, configured to input the text features of the text prompt and the data feature text of the data to be annotated into the large language model, use the large language model to understand the data to be annotated, and generate the data annotation information corresponding to the text prompt for this data to be annotated;

[0040] An algorithm annotation result generation module, configured to, based on the data annotation information, use the data annotation module to complete the data annotation of the data to be annotated.

[0041] A storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above-mentioned data annotation method based on a multimodal large language model are implemented.

[0042] A data annotation device based on a multimodal large language model, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, the steps of the above-mentioned data annotation method based on a multimodal large language model are implemented.

[0043] The beneficial effects of the present invention are as follows: Through the text encoder, the text features of the text prompt can be extracted, and through the feature extraction module, the data features associated with the annotation object in the data to be annotated can be extracted and the data features are mapped into the text space to obtain the data feature text. After the text features and the data feature text are input into the large language model, the large language model understands the data to be annotated and generates the data annotation information, and the data annotation is completed through the data annotation module. In this way, the rapid annotation of the data to be annotated is realized through the large language model;

[0044] In the present invention, based on the judgment conclusion of the accuracy of the algorithm annotation result by the staff, the algorithm annotation result is cleaned based on the accuracy judgment conclusion to obtain a data annotation result, so that the output data annotation result has higher accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 : Flow chart of the present invention.

[0046] Figure 2 : Structural block diagram of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0047] The present invention will be further described in detail below with reference to the drawings and through embodiments. The following embodiments are explanations of the present invention, and the present invention is not limited to the following embodiments.

[0048] Embodiment 1: A data annotation method based on a multimodal large language model for data annotation of image data, including:

[0049] Obtain the text prompt of the data annotation instruction, and extract the text feature of the text prompt through a text encoder;

[0050] Based on the type of data to be annotated, use the corresponding feature extraction module of this type to extract the data features associated with the annotation object in the data to be annotated of this type. In this embodiment, the data to be annotated is image data, and the feature extraction module extracts spatial features from the image data. The feature extraction module for extracting spatial features from image data is an image pre-trained encoder;

[0051] Project the spatial features into the text space to obtain the data feature text corresponding to the spatial features;

[0052] Input the text feature of the text prompt and the data feature text of the image data into the large language model, and use the large language model to understand the image data and generate the data annotation information corresponding to the text prompt for this image data;

[0053] Based on the data annotation information, use the data annotation module to complete the data annotation of the image data.

[0054] In this embodiment, when inputting a text prompt such as "Please mark the position of the face area in the picture with a green rectangle frame and mark the key points such as the left eye, right eye, nose, left corner of the mouth, and right corner of the mouth of the face in the form of coordinate points", the text encoder extracts the text feature of the text prompt. After inputting the text feature into the large language model, the large language model can determine the data annotation instruction by understanding the content of the text feature.

[0055] In this embodiment, after the picture containing face information is input into the feature extraction module, the feature extraction module extracts spatial features from the image data of the picture. The spatial features are mapped to the text space to obtain data feature texts. The data feature texts contain the coordinate position information of the spatial features in the image and can reflect the position information of the spatial features on the picture.

[0056] In this embodiment, after the data feature texts are input into the large language model, the large language model can understand the data feature texts and understand the positions of the spatial features in the picture. The large language model can filter out the position information of the required spatial features according to the data annotation instructions, and perform corresponding analysis on the filtered position information according to the data annotation instructions, so as to obtain data annotation information. The obtained information completes data annotation through the data annotation module to obtain the algorithm annotation result. The large language model annotates the image information of the picture, which greatly improves the annotation speed of the picture image information.

[0057] Furthermore, by showing the algorithm annotation result to the staff and obtaining the accuracy judgment conclusion of the staff on the algorithm annotation result, the algorithm annotation result is cleaned to obtain the data annotation result, so that the output data annotation result has higher accuracy.

[0058] Embodiment 2: A data annotation method based on a multimodal large language model for data annotation of video data, including:

[0059] Obtain the text prompt of the data annotation instruction, and extract the text features of the text prompt through the text encoder;

[0060] Based on the type of the data to be annotated, use the corresponding feature extraction module of this type to extract the data features associated with the annotation object in the data to be annotated of this type. In this embodiment, the data to be annotated is video data, and the feature extraction module extracts spatio-temporal features and language features from the video data. The feature extraction module for extracting spatio-temporal features and language features from the video data is a video pre-training encoder;

[0061] Map the spatial features to the text space to obtain data feature texts corresponding to the spatio-temporal features and language features;

[0062] Input the text features of the text prompt and the data feature texts of the video data into the large language model, and use the large language model to understand the video data and generate the data annotation information of this video data corresponding to the text prompt;

[0063] Based on the data annotation information, use the data annotation module to complete the data annotation of the video data.

[0064] In this embodiment, when a text prompt such as "Please mark the position of the face area in the video with a green rectangular box, and mark key points such as the left eye, right eye, nose, left corner of the mouth, and right corner of the mouth of the face in the form of coordinate points. At the same time, identify the language of the voice in the video. If it is not Chinese, please translate it into Chinese and add Chinese subtitles to the video. Analyze the human emotion from the facial expression, timbre, and pitch of the face" is input, the text encoder extracts the text features of the text prompt. After inputting the text features into the large language model, the large language model can determine the data annotation instruction by understanding the content of the text features.

[0065] In this embodiment, after the video is input into the feature extraction module, the feature extraction module extracts spatio-temporal features and voice features from the video data of the video. The spatio-temporal features and voice features are mapped to the text space, and a data feature text can be obtained. The coordinate position information of the spatial features on each frame of the picture in the video and the language, timbre, and pitch information in the audio can be reflected in the data feature text.

[0066] In this embodiment, after the data feature text is input into the large language model, the large language model can understand the data feature text and understand the position of the spatial features in each frame of the picture in the video and the language, timbre, and pitch in the audio. The large language model can filter out the position information of the required spatial features in each frame of the picture in the video and the language, timbre, and pitch information in the audio according to the data annotation instruction, and perform corresponding analysis on the filtered information to obtain data annotation information. The obtained data annotation information completes data annotation through the data annotation module to obtain an algorithm annotation result. The large language model annotates the image information and audio information of the video, greatly improving the annotation speed of the video information.

[0067] Furthermore, by presenting the algorithm annotation result to the staff and obtaining the accuracy judgment conclusion of the staff on the algorithm annotation result, the algorithm annotation result is cleaned to obtain a data annotation result, making the output data annotation result more accurate.

[0068] Embodiment 3: This embodiment is a data annotation device based on a multi-modal large language model, including:

[0069] A text feature extraction module, configured to obtain a text prompt of a data annotation instruction and extract the text features of the text prompt through a text encoder;

[0070] A data feature extraction module, configured to, based on the type of data to be annotated, use the corresponding feature extraction module of this type to extract the data features associated with the annotation object in the data to be annotated of this type;

[0071] A data feature text acquisition module, configured to map the data features of the data to be annotated into the text space, so as to obtain a data feature text corresponding to the data features of the data to be annotated;

[0072] A data annotation information generation module, configured to input the text features of the text prompt and the data feature text of the data to be annotated into a large language model, use the large language model to understand the data to be annotated, and generate data annotation information corresponding to the text prompt for the data to be annotated;

[0073] An algorithm annotation result generation module, configured to complete the data annotation of the data to be annotated based on the data annotation information by using a data annotation module.

[0074] Embodiment 4: This embodiment is a storage medium, on which a computer program executable by a processor is stored. When the computer program is executed, the steps of the data annotation method based on a multi-modal large language model in Embodiment 1 or 2 are implemented.

[0075] Embodiment 5: This embodiment is a data annotation device based on a multi-modal large language model, which has a memory and a processor. A computer program executable by the processor is stored on the memory. When the computer program is executed, the steps of the data annotation method based on a multi-modal large language model in Embodiment 1 or 2 are implemented.

[0076] Embodiment 6: This embodiment is a human-computer collaborative data annotation and cleaning method based on a multi-modal large language model, including:

[0077] Obtain the text prompt of the data annotation instruction, and extract the text features of the text prompt through a text encoder;

[0078] Based on the type of the data to be annotated, use the corresponding feature extraction module of this type to extract the data features associated with the annotation object in the data to be annotated of this type;

[0079] Map the data features of the data to be annotated into the text space, so as to obtain a data feature text corresponding to the data features of the data to be annotated;

[0080] Input the text features of the text prompt and the data feature text of the data to be annotated into a large language model, use the large language model to understand the data to be annotated, and generate data annotation information corresponding to the text prompt for the data to be annotated;

[0081] Based on the data annotation information, complete the data annotation of the data to be annotated by using a data annotation module to obtain an algorithm annotation result;

[0082] Display the algorithm annotation result to the staff, and obtain the accuracy judgment conclusion of the staff on the algorithm annotation result;

[0083] Based on the accuracy judgment conclusion of the algorithm annotation result, clean the algorithm annotation result to obtain the data annotation result.

[0084] Based on the type of the data to be annotated, use the corresponding feature extraction module of this type to extract the data features associated with the annotation object in the data annotation instruction of this type of data to be annotated, including:

[0085] If the data to be annotated is image data, use the feature extraction module to extract spatial features from the image data;

[0086] If the data to be annotated is video data, use the feature extraction module to extract spatio-temporal features and / or speech features from the video data.

[0087] Embodiment 7: This embodiment is a human-computer collaborative data annotation and cleaning device based on a multi-modal large language model, having:

[0088] A text feature extraction module, used to obtain the text prompt words of the data annotation instruction and extract the text features of the text prompt words through a text encoder;

[0089] A data feature extraction module, used to based on the type of the data to be annotated, use the corresponding feature extraction module of this type to extract the data features associated with the annotation object in the data annotation instruction of this type of data to be annotated;

[0090] A data feature text acquisition module, used to map the data features of the data to be annotated to the text space to obtain the data feature text corresponding to the data features of the data to be annotated;

[0091] A data annotation information generation module, used to input the text features of the text prompt words and the data feature text of the data to be annotated into the large language model, use the large language model to understand the data to be annotated, and generate the data annotation information corresponding to the text prompt words for this data to be annotated;

[0092] An algorithm annotation result generation module, used to based on the data annotation information, use the data annotation module to complete the data annotation of the data to be annotated to obtain the algorithm annotation result;

[0093] An accuracy judgment module, used to display the algorithm annotation result to the staff and obtain the accuracy judgment conclusion of the staff on the algorithm annotation result;

[0094] A data annotation result generation module, used to based on the accuracy judgment conclusion of the algorithm annotation result, clean the algorithm annotation result to obtain the data annotation result.

[0095] Embodiment Eight: This embodiment is a storage medium, on which there is a computer program that can be executed by a processor. When the computer program is executed, it implements the steps of the human-computer collaborative data annotation and cleaning method based on the multimodal large language model in Embodiment Six.

[0096] Embodiment Nine: This embodiment is a data annotation device based on a multimodal large language model, which has a memory and a processor. There is a computer program stored on the memory that can be executed by the processor. When the computer program is executed, it implements the steps of the human-computer collaborative data annotation and cleaning method based on the multimodal large language model in Embodiment Six.

[0097] As mentioned above, the above are only the preferred specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution of the present invention and its inventive concept, makes equivalent substitutions or changes, and all should be covered within the protection scope of the present invention.

Claims

1. A human-machine collaborative data labeling and cleaning method based on a multimodal large language model, characterized in that: include: Obtain text prompt words of the data annotation instruction, and extract text features of the text prompt words through a text encoder; Based on the type of the data to be annotated, a feature extraction module corresponding to the type is used to extract data features associated with the annotated object in the data annotating instruction in the data annotating data of the type; Mapping the data features of the data to be annotated to the text space to obtain the data feature text corresponding to the data features of the data to be annotated; Inputting the text features of the text prompt words and the data feature text of the data to be annotated into the large language model, using the large language model to understand the data to be annotated, and generating data annotation information of the data to be annotated corresponding to the text prompt words; Based on the data annotation information, the data annotation module is used to complete the data annotation of the data to be annotated, and the algorithm annotation result is obtained; Show the algorithm labeling results to the staff and obtain their conclusions on the accuracy of the algorithm labeling results; Based on the accuracy judgment conclusion of the algorithm labeling results, the algorithm labeling results are cleaned to obtain the data labeling results.

2. The method for human-machine collaborative data labeling and cleaning based on a multimodal large language model according to claim 1 is characterized in that: The method of extracting data features associated with the annotation object in the data annotation instruction from the type of the data to be annotated using a feature extraction module corresponding to the type of the data to be annotated includes: If the data to be annotated is image data, a feature extraction module is used to extract spatial features from the image data; If the data to be labeled is video data, a feature extraction module is used to extract spatiotemporal features and / or speech features from the video data.

3. A human-machine collaborative data labeling and cleaning device based on a multimodal large language model, characterized in that: have: A text feature extraction module is used to obtain text prompt words of the data annotation instruction and extract text features of the text prompt words through a text encoder; A data feature extraction module is used to extract data features associated with the annotation object in the data annotation instruction from the type of the data to be annotated using a feature extraction module corresponding to the type; A data feature text acquisition module is used to map the data features of the data to be annotated to the text space to obtain the data feature text corresponding to the data features of the data to be annotated; A data annotation information generation module is used to input the text features of the text prompt word and the data feature text of the data to be annotated into the large language model, use the large language model to understand the data to be annotated, and generate data annotation information of the data to be annotated corresponding to the text prompt word; An algorithm annotation result generation module is used to complete the data annotation of the data to be annotated based on the data annotation information using the data annotation module to obtain the algorithm annotation result; The accuracy judgment module is used to display the algorithm annotation results to the staff and obtain the staff's accuracy judgment conclusion on the algorithm annotation results; The data labeling result generation module is used to judge the conclusion based on the accuracy of the algorithm labeling results, clean the algorithm labeling results, and obtain the data labeling results.

4. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of a human-computer collaborative data labeling and cleaning method based on a multimodal large language model as described in claim 1 or 2 are implemented.

5. A data annotation device based on a multimodal large language model, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps of a human-computer collaborative data labeling and cleaning method based on a multimodal large language model as described in claim 1 or 2 are implemented.

6. A data annotation method based on a multimodal large language model, characterized in that: include: Obtain text prompt words of the data annotation instruction, and extract text features of the text prompt words through a text encoder; Based on the type of the data to be annotated, a feature extraction module corresponding to the type is used to extract data features associated with the annotated object in the data annotating instruction in the data annotating data of the type; Mapping the data features of the data to be annotated to the text space to obtain the data feature text corresponding to the data features of the data to be annotated; Inputting the text features of the text prompt words and the data feature text of the data to be annotated into the large language model, using the large language model to understand the data to be annotated, and generating data annotation information of the data to be annotated corresponding to the text prompt words; Based on the data annotation information, the data annotation module is used to complete the data annotation of the data to be annotated.

7. The data annotation method based on a multimodal large language model according to claim 6 is characterized in that: The method of extracting data features associated with the annotation object in the data annotation instruction from the type of the data to be annotated using a feature extraction module corresponding to the type of the data to be annotated includes: If the data to be annotated is image data, a feature extraction module is used to extract spatial features from the image data; If the data to be labeled is video data, a feature extraction module is used to extract spatiotemporal features and / or speech features from the video data.

8. A data annotation device based on a multimodal large language model, characterized in that: have: A text feature extraction module is used to obtain text prompt words of the data annotation instruction and extract text features of the text prompt words through a text encoder; A data feature extraction module is used to extract data features associated with the annotation object in the data annotation instruction from the type of the data to be annotated using a feature extraction module corresponding to the type; A data feature text acquisition module is used to map the data features of the data to be annotated to the text space to obtain the data feature text corresponding to the data features of the data to be annotated; A data annotation information generation module is used to input the text features of the text prompt word and the data feature text of the data to be annotated into the large language model, use the large language model to understand the data to be annotated, and generate data annotation information of the data to be annotated corresponding to the text prompt word; The algorithm annotation result generation module is used to complete the data annotation of the data to be annotated based on the data annotation information using the data annotation module.

9. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of a data annotation method based on a multimodal large language model described in claim 6 or 7 are implemented.

10. A data annotation device based on a multimodal large language model, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps of a data labeling method based on a multimodal large language model as described in claim 6 or 7 are implemented.

Citation Information

Cited By

  • Video labeling method and device, electronic equipment, storage medium and product

    CN121353994A

  • Man-machine collaborative data annotation method based on large model

    CN122020148A