Image data labeling method and device, computer equipment and readable storage medium

Through the multimodal large model, pre-labeling and auxiliary labeling of image and text information is solved, and the problems of low efficiency and poor consistency of manual labeling in the insurance field are achieved, efficient and accurate image data labeling is achieved, labor costs are reduced and labeling quality is improved.

CN120541253APending Publication Date: 2025-08-26CHINA PING AN PROPERTY INSURANCE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510660246.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-21
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

In the image understanding technology in the insurance field, existing data annotation relies on manual annotation to lead to inefficiency and poor consistency, and there are problems of human error and strong subjectivity.

Method used

A multimodal large model is used to pre-label the image and text information to generate high-quality preliminary labeling results, and provide auxiliary labeling during the manual labeling process. Answers and suggestions are provided through the multimodal labeling model, reducing manual workload and improving labeling accuracy and consistency.

Benefits of technology

Efficient, accurate and consistent image data labeling is achieved, the workload of manual labeling is reduced, the labeling efficiency is improved, and the human misjudgment rate is reduced, ensuring the quality and consistency of the labeling results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120541253A_ABST
    Figure CN120541253A_ABST
Patent Text Reader

Abstract

The invention provides an image data annotation method and device, computer equipment and a readable storage medium, and relates to the field of financial science and technology. The method comprises the steps of obtaining a to-be-labeled image and labeling guide data, wherein the labeling guide data is used for guiding a multi-modal labeling model to label the to-be-labeled image; based on the annotation guide data and the multi-modal annotation model, pre-annotating the to-be-annotated image to obtain a first image description result of the to-be-annotated image; obtaining a second image description result generated by manually annotating the to-be-annotated image based on the first image description result by an annotator; wherein when the annotator carries out manual annotation on the to-be-annotated image, auxiliary annotation is carried out through the multi-modal annotation model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of financial technology, and in particular to an image data annotation method, apparatus, computer equipment and readable storage medium. Background Art

[0002] In the insurance sector, image understanding technology has a wide range of applications, such as: (1) automatically reading and analyzing photos of accident scenes to generate loss assessment reports and speed up claims processing; (2) using on-site images and generated text descriptions to reconstruct the accident process after an accident for liability analysis and insurance claims processing. These applications help improve service quality, reduce costs, and provide a better customer experience.

[0003] Data labeling is a very important part of image understanding technology and determines the upper limit of the model. Data labeling in related technologies mainly relies on manual labeling, which has the following problems: (1) Time-consuming and inefficient: it relies entirely on manual labeling, which is inefficient and slow when processing a large number of images; (2) Manual labeling is highly subjective and has consistency issues: different labelers may have different understandings and labeling standards for the same image, resulting in inconsistent labeling results. Summary of the Invention

[0004] In view of this, the present application provides an image data annotation method, apparatus, computer device and readable storage medium, which can provide an efficient, accurate and reliable annotation solution in image understanding and annotation tasks in the insurance field.

[0005] In a first aspect, an embodiment of the present application provides an image data annotation method, comprising:

[0006] Acquire an image to be annotated and annotation guide data, wherein the annotation guide data is used to guide the multimodal annotation model to annotate the image to be annotated;

[0007] Pre-annotating the image to be annotated based on the annotation guide data and the multimodal annotation model to obtain a first image description result of the image to be annotated;

[0008] Obtain a second image description result generated by an annotator manually annotating the image to be annotated based on the first image description result; wherein, when the annotator manually annotates the image to be annotated, auxiliary annotation is performed through the multimodal annotation model.

[0009] In a second aspect, an embodiment of the present application provides an image data annotation device, comprising:

[0010] A data acquisition module, configured to acquire an image to be annotated and annotation guide data, wherein the annotation guide data is used to guide the multimodal annotation model to annotate the image to be annotated;

[0011] a first labeling module, configured to pre-label the image to be labeled based on the labeling guide data and the multimodal labeling model to obtain a first image description result of the image to be labeled;

[0012] The second annotation module is used to obtain a second image description result generated by the annotator manually annotating the image to be annotated based on the first image description result; wherein, when the annotator manually annotates the image to be annotated, auxiliary annotation is performed through the multimodal annotation model.

[0013] In a third aspect, an embodiment of the present application provides a computer device comprising a processor and a memory, wherein the memory stores programs or instructions that can be run on the processor, and when the programs or instructions are executed by the processor, the steps of the method of the first aspect are implemented.

[0014] In a fourth aspect, an embodiment of the present application provides a readable storage medium, which stores a program or instruction. When the program or instruction is executed by a processor, the steps of the method of the first aspect are implemented.

[0015] The image data annotation method, device, computer equipment and readable storage medium provided in the embodiments of the present application utilize the powerful capabilities of a large multimodal model, combine image and text information, generate high-quality preliminary annotations, avoid the introduction of human errors, and ensure the accuracy and consistency of the annotation results. Moreover, pre-annotation through a multimodal annotation model greatly reduces the workload of manual annotation, saves time, and improves annotation efficiency. In addition, in the process of the annotator manually annotating the image to be annotated based on the first image description result, if an uncertain problem is encountered, the multimodal annotation model can be used for auxiliary annotation, and the multimodal annotation model can provide answers and suggestions, which further improves the accuracy and consistency of the annotation, reduces the human error rate, and can minimize the labor cost.

[0016] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present application are listed below. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0018] Figure 1 A schematic diagram showing an application environment of the image data annotation method according to an embodiment of the present application is shown;

[0019] Figure 2 FIG1 shows one of the flow charts of the image data annotation method according to an embodiment of the present application;

[0020] Figure 3 The second flowchart of the image data annotation method according to the embodiment of the present application is shown;

[0021] Figure 4 A structural block diagram of an image data annotation device according to an embodiment of the present application is shown;

[0022] Figure 5 One of the structural block diagrams of the computer device according to an embodiment of the present application is shown;

[0023] Figure 6 The second structural block diagram of the computer device according to the embodiment of the present application is shown. DETAILED DESCRIPTION

[0024] The following will be combined with the accompanying drawings in the embodiments of the present application to clearly describe the technical solutions in the embodiments of the present application. Obviously, the embodiments described are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field are within the scope of protection of this application.

[0025] The terms "first," "second," and the like in the specification and claims of this application are used to distinguish similar objects, and are not used to describe a specific order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of this application can be implemented in an order other than that illustrated or described herein, and that the objects distinguished by "first," "second," and the like are generally of the same type, and do not limit the number of objects; for example, the first object can be one or more. In addition, the term "and / or" in the specification and claims refers to at least one of the connected objects, and the character " / " generally indicates that the objects connected are in an "or" relationship.

[0026] The image data annotation method based on the multimodal large model provided in the embodiment of the present application can be applied in the following fields: Figure 1 In an application environment, a client communicates with a server via a network. The client can obtain an image to be annotated and annotation guide data, the annotation guide data being used to guide a multimodal annotation model in annotating the image to be annotated; based on the annotation guide data and the multimodal annotation model, the image to be annotated is pre-annotated to obtain a first image description result of the image to be annotated; and a second image description result generated by an annotator manually annotating the image to be annotated based on the first image description result is obtained; wherein, when the annotator manually annotates the image to be annotated, the multimodal annotation model is used to assist in the annotation.

[0027] The embodiment of the present application utilizes the powerful capabilities of a large multimodal model, combines image and text information, generates high-quality preliminary annotations, avoids the introduction of human errors, and can ensure the accuracy and consistency of the annotation results. Moreover, by pre-annotating with a multimodal annotation model, the workload of manual annotation is greatly reduced, time is saved, and annotation efficiency is improved. In addition, in the process of the annotator manually annotating the image to be annotated based on the first image description result, if an uncertain problem is encountered, the multimodal annotation model can be used for auxiliary annotation, and the multimodal annotation model can provide answers and suggestions, which further improves the accuracy and consistency of the annotation, reduces the human error rate, and can minimize the labor cost.

[0028] The client can be, but is not limited to, various personal computers, laptops, smart phones, tablet computers, and portable wearable devices. The server can be implemented as an independent server or a server cluster consisting of multiple servers.

[0029] The image data annotation method based on a multimodal large model provided in the embodiments of this application can be applied in the field of financial technology, specifically in insurance scenarios, such as auto insurance, health insurance, property insurance, etc. Below, in conjunction with the accompanying drawings, the image data annotation method, device, computer device, and readable storage medium provided in the embodiments of this application are described in detail through specific embodiments and their application scenarios. The following embodiments and features of the embodiments can be combined with each other unless they conflict.

[0030] The present application embodiment provides an image data annotation method, such as Figure 2 As shown, the method includes:

[0031] Step 201 : Acquire an image to be annotated and annotation guide data, where the annotation guide data is used to guide a multimodal annotation model to annotate the image to be annotated.

[0032] In this step, we obtain the image to be annotated and the annotation guide data, and store the image to be annotated in a MySQL database. The annotation guide data explains the annotation task, labeling system, and precautions in detail, providing guidance for the multimodal annotation model to perform the annotation work.

[0033] Step 202 : pre-annotate the image to be annotated based on the annotation guide data and the multimodal annotation model to obtain a first image description result of the image to be annotated.

[0034] In this step, the annotation guide data and the image to be annotated are input into the multimodal annotation model, and a first image description result of the image to be annotated is output, where the first image description result is text data.

[0035] For example, take the "image description" annotation task as an example:

[0036] Input 1: Annotation guide, for example, please describe in detail the scene and what happened in the picture;

[0037] Input 2: image to be annotated, for example, 1.jpg;

[0038] Annotation process: The annotation guide and the image to be annotated are input into the multimodal large model to generate the annotation results, that is, the image description, such as the vehicle's left front light is broken and the bumper is detached in the image.

[0039] The multimodal annotation model is a large multimodal model, obtained through pre-training. Taking the automobile insurance loss assessment scenario as an example, a multimodal claims dataset is constructed using annotated accident scene photos and a historical claims case database. This dataset is then preprocessed, for example, using YOLOv8 to detect license plates and vehicle frame numbers, automatically associating them with policy information, and enhancing the model's robustness through image enhancement (e.g., simulating nighttime lighting and rain and snow obstructions). A deep learning model is then trained based on this dataset to produce the multimodal annotation model.

[0040] The present embodiment leverages the powerful capabilities of a large multimodal model, combining image and text information to generate high-quality preliminary annotations, avoiding the introduction of human error and ensuring the accuracy and consistency of the annotation results. Furthermore, pre-annotation using a multimodal annotation model significantly reduces the workload of manual annotation, saving time and improving annotation efficiency.

[0041] Step 203 , obtaining a second image description result generated by the annotator manually annotating the image to be annotated based on the first image description result; wherein, when the annotator manually annotates the image to be annotated, auxiliary annotation is performed using the multimodal annotation model.

[0042] In this step, the annotator manually annotates the image to be annotated, referring to the first image description result, to obtain a second image description result. While the annotator is manually annotating the image to be annotated, real-time interaction is carried out, and assisted annotation is achieved through the multimodal annotation model, whose assisted annotation function is pre-established through model training.

[0043] In one embodiment of the present application, when an annotator manually annotates an image to be annotated, auxiliary annotation is performed using a multimodal annotation model, including:

[0044] Obtaining annotation problem information of the annotator for the first image description result;

[0045] The annotation problem information is input into the multimodal annotation model, and feedback information regarding the annotation problem information is output. The feedback information is used to guide the annotator to manually annotate the image to be annotated.

[0046] In this embodiment, the annotator raises a question regarding the first image description result. This question is fed into the multimodal annotation model, which then outputs feedback regarding the question. The annotator then manually annotates the image to be annotated based on this feedback to generate a second image description result. The assisted annotation function can be initiated by the annotator, who enters a voice / text question via a pop-up window. Alternatively, it can be passively triggered, monitoring the annotator's annotation behavior. When hesitation (e.g., repeated adjustments to the annotation box) or anomalies (e.g., the annotation position conflicts with historical data) are detected, a pop-up window automatically suggests using the assisted annotation function.

[0047] For example, in an auto insurance claim annotation application, the first image description results in a deformation on a side view of the vehicle. The annotator is unsure whether this location belongs to a "door" or a "side skirt," and asks a question. The multimodal annotation model analyzes geometric features in the image, such as the correlation between the deformed area and the door gap. Combined with historical annotation data, it determines that the historical damage to the side skirts of this vehicle model accounts for only 3%. It then responds with a "door deformation" and a confidence level, such as "85% confidence recommended, manual review recommended." Based on this response, the annotator can label the area "door deformation" or, if they believe the model's answer is incorrect, "side skirt deformation."

[0048] In a car insurance claims annotation application, consider damage type determination: the first image description indicates a crack in the vehicle's glass. The annotator asks whether it is a "star-shaped crack" or a "linear crack." The multimodal annotation model identifies the crack morphology through image segmentation and suggests the label "star-shaped crack." Based on this answer, the annotator can label it "glass star-shaped crack" or, if they believe the model's answer is incorrect, "glass linear crack."

[0049] In the embodiment of the present application, when the annotator manually annotates the image to be annotated based on the first image description result, if an uncertain problem is encountered, the multimodal annotation model can be used for auxiliary annotation. The multimodal annotation model provides answers and suggestions, further improving the accuracy and consistency of the annotation, reducing the human error rate, and at the same time minimizing the labor cost.

[0050] In one embodiment of the present application, auxiliary annotation can be used to score the annotation skills of annotators. For example, the number of times an annotator uses auxiliary annotation can be counted, and the annotator's annotation skills can be scored based on this number. For example, the more times an annotator uses auxiliary annotation, the lower the annotator's annotation skills score. For annotators with low scores, subsequent professional training can be provided to improve their annotation skills and ensure the quality of image annotation.

[0051] In one embodiment of the present application, the method further includes: comparing the first image description result with the second image description result, and issuing a reminder message if the similarity between the first image description result and the second image description result is less than or equal to a preset similarity threshold.

[0052] In this embodiment, the annotator obtains a second image description result by self-annotation without using the multimodal annotation model for auxiliary annotation, or obtains a second image description result by using the multimodal annotation model for auxiliary annotation. After obtaining the second image description result, the second image description result is compared with the first image description result. If the similarity between the two is less than or equal to a preset similarity threshold, indicating that the model annotation and the manual annotation are inconsistent, a reminder message is issued to remind the annotator to check the annotation results.

[0053] In the related art, the quality control of annotation is insufficient. The quality control in the related art usually relies on post-audit and lacks real-time annotation quality monitoring and feedback, making it difficult to correct errors in a timely manner. Considering this problem, the embodiment of the present invention provides another image data annotation method, such as Figure 3 As shown, the method includes:

[0054] Step 301: Acquire an image to be annotated and annotation guide data, where the annotation guide data is used to guide the multimodal annotation model to annotate the image to be annotated.

[0055] Step 302: pre-annotate the image to be annotated based on the annotation guide data and the multimodal annotation model to obtain a first image description result of the image to be annotated;

[0056] Step 303: Obtain a second image description result generated by the annotator manually annotating the image to be annotated based on the first image description result; wherein, when the annotator manually annotates the image to be annotated, auxiliary annotation is performed using the multimodal annotation model;

[0057] Step 304: Evaluate the first image description result using multiple dimensional indicators to obtain a score corresponding to each dimensional indicator; wherein the multiple dimensional indicators include at least two indicators of completeness, toxicity, and redundancy;

[0058] Step 305 : Multiply each scoring score by the weight coefficient of its corresponding indicator and then add them together to obtain a quality inspection score of the first image description result.

[0059] Steps 301 to 303 are the same or similar to steps 201 to 203 in the above embodiment and are not described in detail here. Steps 304 and 305 can be performed after or before step 303.

[0060] In this embodiment, the first image description result is evaluated based on multiple metrics to obtain a quality inspection score for the first image description result. Specifically, the first image description result is evaluated based on at least two metrics of integrity, toxicity, and redundancy to obtain scores corresponding to each metric. The scores are then weighted and summed to obtain a quality inspection score for the first image description result.

[0061] In one embodiment of the present application, the first image description result is evaluated for multiple dimensional indicators to obtain a score corresponding to each dimensional indicator, including:

[0062] For completeness assessment, the object detection model is used to identify key objects appearing in the image to be annotated, and the large language model is used to detect the description information of the key objects in the first image description result. The identified key objects are compared with the description information of the key objects to determine the completeness score of the first image description result;

[0063] For the evaluation of toxicity, the first image description result is evaluated by a toxicity scoring model, and a toxicity score of the first image description result is output;

[0064] For redundancy evaluation, the first image description result is evaluated using a text redundancy detection model, and a redundancy score of the first image description result is output.

[0065] In this embodiment, an object detection model is used to detect key objects appearing in the image to be annotated, and an LLM (Large Language Model) is used to detect the description information of the key objects in the first image description result. The identified key objects are compared with the description information of the key objects to determine whether the first image description result completely describes the information of the key objects, and a score for the completeness of the first image description result is obtained, thereby avoiding missing key objects.

[0066] The toxicity of a description includes whether it contains insulting, sensitive, pornographic, or political content. This can be assessed using an open-source toxicity scoring model. This model accepts text input and outputs scores for each type of toxicity within the description, along with a final toxicity score. The more toxicity, the lower the score. For example, given an image of a torn car bumper, the outputs are: insult score 10; sensitivity score 10; pornography score 10; political score 10; and a final toxicity score of 10.

[0067] Pre-train a text redundancy detection model, for example, using the BERT model as a base. Use this text redundancy detection model to evaluate the first image description result and obtain a redundancy score for the first image description result. The more redundancy in the text, the lower the score. For example, the text "Hello, hello, hello, hello" has a score of 0; the text "Hello, I am an AI model" has a score of 10.

[0068] The embodiment of the present application can perform quality inspection and scoring of the pre-annotation results to achieve real-time annotation quality monitoring.

[0069] Furthermore, in one embodiment of the present application, the method also includes: if the quality inspection score is greater than or equal to a preset threshold, determining that the first image description result has passed the quality inspection; if the quality inspection score is less than the preset threshold, triggering the pre-labeling process and / or manual labeling process.

[0070] In this embodiment, the quality inspection score is compared with a preset threshold to determine the quality inspection result. If the quality inspection score is greater than or equal to the preset threshold, it indicates that the first image description result meets the indicator requirements, and the first image description result is determined to have passed the quality inspection, and the passed quality inspection data is stored in the database. If the quality inspection score is less than the preset threshold, it indicates that the first image description result does not meet the indicator requirements, and the pre-labeling process and / or manual labeling process are triggered to perform re-labeling.

[0071] In the embodiment of the present application, quality control can be performed after quality inspection and scoring to ensure that the labeling results meet preset standards and ensure that the labeling results are accurate and harmless.

[0072] In one embodiment of the present application, the method further includes: adjusting the weight coefficient corresponding to each dimensional indicator based on business needs.

[0073] In this embodiment, the weight coefficients corresponding to the indicators of each dimension are adjusted according to business needs. For example, for the task of object detection, which places greater emphasis on completeness, the weight coefficient of the completeness indicator is higher than the weight coefficients of toxicity and redundancy; for the task of image description, which places greater emphasis on toxicity, the weight coefficient of the toxicity indicator is higher than the weight coefficients of completeness and redundancy.

[0074] In the embodiment of the present application, the weight coefficients corresponding to the indicators are adjusted so that the quality assessment is more adapted to business needs and the quality assessment is more accurate.

[0075] As a specific implementation of the above-mentioned image data annotation method, the embodiment of the present application provides an image data annotation device. Figure 4 As shown, the image data annotation device 400 includes: a data acquisition module 401 , a first annotation module 402 and a second annotation module 403 .

[0076] The data acquisition module 401 is used to acquire the image to be annotated and the annotation guide data, and the annotation guide data is used to guide the multimodal annotation model to annotate the image to be annotated;

[0077] A first annotation module 402 is configured to pre-annotate the image to be annotated based on the annotation guide data and the multimodal annotation model to obtain a first image description result of the image to be annotated;

[0078] The second annotation module 403 is used to obtain a second image description result generated by the annotator manually annotating the image to be annotated based on the first image description result; wherein, when the annotator manually annotates the image to be annotated, auxiliary annotation is performed through the multimodal annotation model.

[0079] Furthermore, the first labeling module 402 is further configured to:

[0080] Obtaining annotation problem information of the annotator for the first image description result;

[0081] The annotation problem information is input into the multimodal annotation model, and feedback information regarding the annotation problem information is output. The feedback information is used to guide the annotator to manually annotate the image to be annotated.

[0082] Furthermore, the device further includes: a comparison module, configured to:

[0083] The first image description result and the second image description result are compared, and if the similarity between the first image description result and the second image description result is less than or equal to a preset similarity threshold, a reminder message is issued.

[0084] Furthermore, the device further includes: an evaluation module, configured to:

[0085] Evaluate the first image description result using multiple dimensional indicators to obtain a score corresponding to each dimensional indicator;

[0086] Multiply each score by the weight coefficient of its corresponding indicator and add them together to obtain the quality inspection score of the first image description result;

[0087] The multi-dimensional indicators include at least two indicators of integrity, toxicity and redundancy.

[0088] Furthermore, the evaluation module is specifically used to:

[0089] For completeness assessment, the object detection model is used to identify key objects appearing in the image to be annotated, and the large language model is used to detect the description information of the key objects in the first image description result. The identified key objects are compared with the description information of the key objects to determine the completeness score of the first image description result;

[0090] For the evaluation of toxicity, the first image description result is evaluated by a toxicity scoring model, and a toxicity score of the first image description result is output;

[0091] For redundancy evaluation, the first image description result is evaluated using a text redundancy detection model, and a redundancy score of the first image description result is output.

[0092] Furthermore, the evaluation module is also used to:

[0093] If the quality inspection score is greater than or equal to the preset threshold, it is determined that the first image description result has passed the quality inspection;

[0094] If the quality inspection score is less than the preset threshold, the pre-labeling process and / or manual labeling process is triggered.

[0095] Furthermore, the evaluation module is also used to:

[0096] Based on business needs, adjust the weight coefficient corresponding to each dimension indicator.

[0097] The image data annotation device 400 in the embodiment of the present application can be a computer device, or a component in a computer device, such as an integrated circuit or a chip. The computer device can be a terminal, or a device other than a terminal. For example, the computer device can be a mobile phone, a tablet computer, a laptop computer, a PDA, a mobile Internet device (MID), a robot, an ultra-mobile personal computer (UMPC), a netbook or a personal digital assistant (PDA), etc. It can also be a server, a network attached storage (NAS), a personal computer (PC), etc., and the embodiment of the present application does not specifically limit it.

[0098] The image data annotation device 400 provided in the embodiment of the present application can achieve Figure 2 and Figure 3 To avoid repetition, the various processes implemented in the embodiment of the image data annotation method are not described here.

[0099] The embodiment of the present application further provides a computer device, which may be a client, and its internal structure diagram may be as shown in FIG. Figure 5As shown. The computer device includes a processor, memory, network interface, display screen and input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it implements the functions or steps of the client side of the image data annotation method.

[0100] The present application also provides a computer device, such as Figure 6 As shown, the computer device 600 includes a processor 601 and a memory 602. The memory 602 stores programs or instructions that can be run on the processor 601. When the program or instructions are executed by the processor 601, the various steps of the above-mentioned image data annotation method embodiment are implemented and the same technical effect can be achieved. To avoid repetition, they are not repeated here.

[0101] The memory 602 can be used to store software programs and various data. The memory 602 may mainly include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area may store an operating system, applications or instructions required for at least one function (such as a sound playback function, an image playback function, etc.). In addition, the memory 602 may include volatile memory or non-volatile memory, or the memory 602 may include both volatile and non-volatile memory. The non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct RAM bus random access memory (DRRAM). The memory 602 in the embodiment of the present application includes but is not limited to these and any other suitable types of memory.

[0102] Processor 601 may include one or more processing units. Optionally, processor 601 integrates an application processor and a modem processor. The application processor primarily handles operations related to the operating system, user interface, and application programs, while the modem processor primarily processes wireless communication signals, such as a baseband processor. It is understood that the modem processor may not be integrated into processor 601.

[0103] An embodiment of the present application also provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the various processes of the above-mentioned image data annotation method embodiment are implemented and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.

[0104] An embodiment of the present application also provides a chip, which includes a processor and a communication interface. The communication interface and the processor are coupled. The processor is used to run programs or instructions to implement the various processes of the above-mentioned image data annotation method embodiment and can achieve the same technical effect. To avoid repetition, it will not be repeated here.

[0105] It should be understood that the chip mentioned in the embodiments of the present application can also be called a system-level chip, a system chip, a chip system or a system-on-chip chip, etc.

[0106] An embodiment of the present application also provides a computer program product, which is stored in a storage medium and is executed by at least one processor to implement the various processes of the above-mentioned image data annotation method embodiment and can achieve the same technical effect. To avoid repetition, it will not be repeated here.

[0107] It should be noted that, in this article, the terms "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the statement "comprises a ..." does not exclude the presence of other identical elements in the process, method, article or device comprising the element. In addition, it should be noted that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in the opposite order according to the functions involved. For example, the described method may be performed in an order different from that described, and various steps may also be added, omitted, or combined. In addition, the features described with reference to certain examples may be combined in other examples.

[0108] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of this application, ordinary technicians in this field can also make many forms without departing from the purpose of this application and the scope of protection of the claims, all of which are within the protection of this application.

Claims

1. A method for labeling image data, characterized in that: include: Acquire an image to be annotated and annotation guide data, wherein the annotation guide data is used to guide the multimodal annotation model to annotate the image to be annotated; Pre-annotating the image to be annotated based on the annotation guide data and the multimodal annotation model to obtain a first image description result of the image to be annotated; Obtain a second image description result generated by an annotator manually annotating the image to be annotated based on the first image description result; wherein, when the annotator manually annotates the image to be annotated, auxiliary annotation is performed through the multimodal annotation model.

2. The method according to claim 1, characterized in that When the annotator manually annotates the image to be annotated, auxiliary annotation is performed by using the multimodal annotation model, including: Obtaining annotation problem information of the annotator for the first image description result; The labeling problem information is input into the multimodal labeling model, and feedback information for the labeling problem information is output, where the feedback information is used to guide the annotator to manually label the image to be labeled.

3. The method according to claim 1, characterized in that The method further comprises: The first image description result and the second image description result are compared, and if the similarity between the first image description result and the second image description result is less than or equal to a preset similarity threshold, a reminder message is issued.

4. The method according to claim 1, wherein The method further comprises: Evaluate the first image description result based on multiple dimensional indicators to obtain a score corresponding to each dimensional indicator; Multiplying each of the scoring scores by the weight coefficient of its corresponding indicator and then adding them up to obtain a quality inspection score of the first image description result; The multi-dimensional indicators include at least two indicators of integrity, toxicity and redundancy.

5. The method according to claim 4, characterized in that The first image description result is evaluated by multi-dimensional indicators to obtain a score corresponding to each dimensional indicator, including: For the completeness assessment, key objects appearing in the image to be annotated are identified using an object detection model, and description information of the key objects in the first image description result is detected using a large language model, and the identified key objects are compared with the description information of the key objects to determine a score for the completeness of the first image description result; For the evaluation of the toxicity, the first image description result is evaluated using a toxicity scoring model, and a toxicity score of the first image description result is output; For the redundancy evaluation, the first image description result is evaluated by a text redundancy detection model, and a redundancy score of the first image description result is output.

6. The method according to claim 4, characterized in that The method further comprises: If the quality inspection score is greater than or equal to a preset threshold, determining that the first image description result passes the quality inspection; If the quality inspection score is less than a preset threshold, the pre-labeling process and / or the manual labeling process is triggered.

7. The method according to claim 4, characterized in that The method further comprises: Based on business needs, adjust the weight coefficient corresponding to each dimension indicator.

8. An image data annotation device, characterized in that: include: A data acquisition module, configured to acquire an image to be annotated and annotation guide data, wherein the annotation guide data is used to guide the multimodal annotation model to annotate the image to be annotated; a first labeling module, configured to pre-label the image to be labeled based on the labeling guide data and the multimodal labeling model to obtain a first image description result of the image to be labeled; The second annotation module is used to obtain a second image description result generated by the annotator manually annotating the image to be annotated based on the first image description result; wherein, when the annotator manually annotates the image to be annotated, auxiliary annotation is performed through the multimodal annotation model.

9. A computer device, characterized in that: The method comprises a processor and a memory, wherein the memory stores a program or instruction running on the processor, and when the program or instruction is executed by the processor, the steps of the image data annotation method according to any one of claims 1 to 7 are implemented.

10. A readable storage medium having a program or instruction stored thereon, characterized in that: When the program or instruction is executed by a processor, the steps of the image data annotation method according to any one of claims 1 to 7 are implemented.