Image labeling method and device, electronic equipment and storage medium

By using interactive strategies and annotation models to automate image data processing, the problem of low efficiency in manual annotation is solved, and efficient road data annotation is achieved.

CN117197514BActive Publication Date: 2026-05-12Z-ONE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Z-ONE TECH CO LTD
Filing Date
2022-05-31
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing technologies for manually labeling road data are inefficient and struggle to efficiently handle the demand for labeling massive amounts of high-quality road data.

Method used

The image to be labeled is processed by a predetermined interaction strategy to obtain interaction information, which is then input into a pre-trained labeling model to generate a label file indicating the type of the target object. The labeling process does not require human intervention.

Benefits of technology

It has achieved efficient and automated road data annotation, improved annotation efficiency, and solved the problem of time-consuming manual annotation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117197514B_ABST
    Figure CN117197514B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide an image labeling method and device, electronic equipment and storage medium, the method comprising: obtaining a to-be-labeled image; processing the to-be-labeled image through a pre-determined interaction strategy to obtain interaction information, wherein the interaction information is used to indicate the position of a target object in the to-be-labeled image; inputting the interaction information and the to-be-labeled image into a pre-trained labeling model to obtain a labeling file output by the labeling model, wherein the labeling file output by the labeling model is used to indicate the type of the target object in the to-be-labeled image. By applying the present application, the problem of long time consumption of manual labeling of road data and low efficiency of labeling road data can be solved, and the efficiency of image labeling is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and in particular to an image annotation method, apparatus, electronic device, and storage medium. Background Technology

[0002] With the development of technology, autonomous driving is becoming more and more widespread. In order for autonomous driving algorithms to handle more and more complex scenarios, a large amount of real road data is needed to support them. Therefore, the demand for road data annotation is getting higher and higher. Massive and high-quality, detailed road annotation data can greatly improve the safety and practicality of autonomous driving.

[0003] Currently, road data is labeled manually, such as by labeling cars, roads, people, buildings, and vegetation in images. However, manual labeling of road data is time-consuming, resulting in low efficiency.

[0004] Therefore, how to efficiently annotate massive amounts of road data has become an urgent problem to be solved. Summary of the Invention

[0005] In view of this, embodiments of this application provide an image annotation method, apparatus, electronic device, and storage medium to at least partially solve the above-mentioned problems.

[0006] According to a first aspect of the embodiments of this application, an image annotation method is provided, the method comprising: acquiring an image to be annotated; processing the image to be annotated through a predetermined interaction strategy to obtain interaction information, wherein the interaction information is used to indicate the position of a target object in the image to be annotated; inputting the interaction information and the image to be annotated into a pre-trained annotation model to obtain an annotation file output by the annotation model, wherein the annotation file output by the annotation model is used to indicate the type of the target object in the image to be annotated.

[0007] According to a second aspect of the embodiments of this application, an image annotation apparatus is provided, the apparatus comprising: an acquisition module for acquiring an image to be annotated; an image processing module for processing the image to be annotated through a predetermined interaction strategy to obtain interaction information, wherein the interaction information is used to indicate the position of a target object in the image to be annotated; and an annotation module for inputting the interaction information and the image to be annotated into a pre-trained annotation model to obtain an annotation file output by the annotation model, wherein the annotation file output by the annotation model is used to indicate the type of the target object in the image to be annotated.

[0008] According to a third aspect of the present application, an electronic device is provided, comprising: a processor, a memory, a communication interface, and a communication bus, wherein the processor, the memory, and the communication interface communicate with each other via the communication bus; the memory is used to store at least one executable instruction, wherein the executable instruction causes the processor to perform an operation corresponding to the method described in the first aspect.

[0009] According to a fourth aspect of the embodiments of this application, a computer storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the method described in the first aspect.

[0010] Based on the annotation method provided by the above scheme, the image to be annotated is processed through a pre-determined interaction strategy to obtain interaction information. This interaction information, along with the image to be annotated, is then input into a pre-trained annotation model. Through inference by the annotation model, an annotation file indicating the type of target object in the image can be obtained. It is evident that interaction information can be obtained through the interaction strategy, indicating the position of objects in the image. The annotation model quickly identifies the position of target objects in the image based on this interaction information, thereby identifying the type of target objects and generating an annotation file indicating the type of target objects. The annotation process requires no manual intervention, solving the problems of time-consuming and inefficient manual annotation, and achieving efficient annotation of massive amounts of road data. Attached Figure Description

[0011] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings.

[0012] Figure 1 A flowchart illustrating an image annotation method provided in one embodiment of this application;

[0013] Figure 2 A flowchart illustrating a street view data annotation method provided in one embodiment of this application;

[0014] Figure 3 A flowchart of a labeled model training method provided in one embodiment of this application;

[0015] Figure 4 This is a schematic diagram of an image annotation apparatus provided in one embodiment of this application;

[0016] Figure 5 This is a schematic diagram of the structure of an electronic device provided in one embodiment of this application. Detailed Implementation

[0017] To enable those skilled in the art to better understand the technical solutions in the embodiments of this application, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art should fall within the protection scope of the embodiments of this application.

[0018] The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. The singular forms “a,” “the,” and “the” used in this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.

[0019] It should be understood that although the terms first, second, third, etc., may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, without departing from the scope of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to determination."

[0020] like Figure 1 As shown, Figure 1 This is a flowchart of an image annotation method provided in an embodiment of this application, which includes the following steps 101 to 103:

[0021] Step 101: Obtain the image to be labeled.

[0022] Retrieve the image to be labeled from a local database or a cloud database. The image to be labeled can be a picture or a video frame, and can be of any type, such as: landscape pictures, portrait pictures, road videos, etc.

[0023] Step 102: Process the image to be labeled using a predetermined interaction strategy to obtain interaction information, wherein the interaction information is used to indicate the position of the target object in the image to be labeled.

[0024] After acquiring the image to be labeled, it is processed through a pre-determined interaction strategy. Once the processing is complete, interaction information can be obtained, including the position of the labeled object, its edges, and the positive and negative points of the labeling results, etc.

[0025] Step 103: Input the interaction information and the image to be labeled into the pre-trained labeling model to obtain the labeling file output by the labeling model. The labeling file output by the labeling model is used to indicate the type of target object in the image to be labeled.

[0026] A pre-trained annotation model is used to label the types of target objects in an input image, generating a label file indicating the types of target objects. The annotation model can identify one type of target object or multiple types. When the annotation model identifies one type of target object, the output label file indicates whether the image to be annotated contains the corresponding type of target object. When the annotation model identifies multiple types of target objects, the output label file indicates the type of each target object in the image to be annotated.

[0027] After acquiring the interaction information, the interaction information and the image to be labeled are input into the pre-trained labeling model. The labeling model can obtain the label file through inference. The label file can indicate the type of target object in the image to be labeled. The target object can be a flower, a person, a car, a street lamp, etc.

[0028] In this embodiment, the image to be labeled is processed using a pre-determined interaction strategy to obtain interaction information. This interaction information, along with the image to be labeled, is then input into a pre-trained labeling model. Through inference by the labeling model, a labeled file indicating the type of target object in the image can be obtained. Thus, the interaction strategy yields interaction information, which indicates the position of an object in the image. Based on this interaction information, the labeling model quickly identifies the position of the target object in the image, thereby identifying the type of the target object and generating a labeled file indicating that type. The labeling process requires no manual intervention, solving the problems of time-consuming and inefficient manual labeling, and achieving efficient labeling of massive amounts of road data.

[0029] In one possible implementation, before processing the image to be labeled using a pre-determined interaction strategy, the image labeling task can be obtained, and then the interaction strategy and labeling model can be determined based on the image labeling task. Different image labeling tasks correspond to different combinations of interaction strategies and labeling models.

[0030] Different application scenarios present different annotation tasks, such as: annotating only people in a person image, annotating other organisms in a person image, or annotating only plants, etc. Different annotation tasks correspond to different annotation strategies and annotation model combinations. For example, if the target object is a person, the first interaction strategy and the first annotation model can be used; if the target object is a plant, the second interaction strategy and the second annotation model can be used.

[0031] It should be understood that multiple interaction strategies and annotation models are pre-created. Depending on the annotation task, the interaction strategies and annotation models can be freely combined to create an interaction strategy and annotation model suitable for the corresponding annotation task.

[0032] In this embodiment of the application, after determining the annotation task, the corresponding interaction strategy and annotation model can be determined according to the annotation task. Different interaction strategies are suitable for identifying the regions where different types of target objects are located in the image to be identified, and different annotation models are suitable for identifying different types of target objects. By determining the interaction strategy and annotation model according to the annotation task, different types of target objects can be annotated, thus ensuring the applicability of the image annotation method.

[0033] In one possible implementation, when determining the interaction strategy based on the image annotation task, at least one target object to be annotated can be determined based on the image annotation task, and then an interaction strategy matching the type of the target object can be determined from at least two preset interaction strategies based on the type of the target object.

[0034] Multiple interaction strategies are pre-created for different interaction tasks, with different strategies corresponding to different tasks. Each interaction strategy defines a method for identifying the region containing a target object in an image. After obtaining the annotation task, the interaction strategy corresponding to the annotation task can be determined from the pre-created strategies. For example, when the interaction task is to perform high-precision identification of a target object, the interaction strategy is a seed-based identification method; when the interaction task is to identify and detect target types, the interaction strategy is a region of interest (ROI)-based identification method.

[0035] In this embodiment of the application, by selecting an interaction strategy corresponding to different interaction tasks, it is possible to quickly and accurately identify the regions where different types of target objects are located in the image to be identified, and generate interaction information containing the location information of the target objects. Thus, the annotation model can quickly identify the location of the target objects in the image to be identified based on the interaction information, thereby improving the efficiency of annotation.

[0036] Optionally, point-based recognition methods include boundary seed recognition methods and region seed recognition methods. Boundary seed recognition methods are used to recognize target objects with clear boundaries, while region seed recognition methods are used to recognize target objects with unclear boundaries.

[0037] Optionally, the region of interest (ROI) based identification methods include drawn ROI identification, tight ROI BB identification, and loose ROI BB identification. The drawn ROI identification method is used to identify target objects with clear boundaries, the tight ROI identification method is used to identify target objects with unclear boundaries and multiple targets, and the loose ROI identification method is used to identify target objects with unclear boundaries and only one target.

[0038] In one possible implementation, before inputting the interactive information and the image to be labeled into the pre-trained labeling model, a street view dataset including street view data can be obtained. The pre-built model can then be trained using the street view dataset to obtain the labeling model. The labeling model can then be deployed to the cloud, and the cloud-deployed labeling model can be encapsulated to obtain a web interface for calling the labeling model.

[0039] The dataset can be a publicly available street view dataset or a private street view dataset collected by a data collection vehicle. The acquired dataset is used to train a pre-built labeled model. The trained model has the ability to identify the type of target object. The trained model is then deployed to the cloud and encapsulated in the cloud into an interface that can be input and output on the web to prepare for subsequent policy interaction and model inference.

[0040] In this embodiment of the application, a labeling model is trained, and then the trained labeling model is deployed to the cloud and packaged. By training the labeling model, the ability of the labeling model to identify target objects can be improved, and the accuracy of labeling can be improved. Deploying the labeling model to the cloud can be done without being restricted by the location or device used, supports multi-person collaborative work, and provides overall control from raw data to labeling results, thereby improving labeling efficiency.

[0041] In one possible implementation, after obtaining the annotation file output by the annotation model, it can be checked whether the first annotation label in the annotation file, which indicates the type of the target object, matches at least one preset second annotation label; if the first annotation label matches at least one second annotation label, the annotation file is output.

[0042] The first label is the annotation label assigned to the annotation file by the annotation model, and the second label is a pre-set annotation label. Different annotation tasks correspond to different pre-set labels. For example, if the annotation task is to detect target types, the corresponding target categories need to be pre-set; if the task is to identify target colors, different color categories need to be pre-set, and so on. During the annotation process of the annotation model on the image to be annotated, the annotation model assigns an annotation label to the annotation file based on the annotation results. This annotation label may include the target type, color, size, etc., depending on the annotation task. The annotation label assigned by the annotation model is matched with the pre-configured annotation labels. If they match, it means that the annotation result is reliable, and then the annotation file is output.

[0043] In one possible implementation, if the first annotation label does not match at least one second annotation label, the image to be annotated, the interaction information, and the annotation file are input into the annotation model to re-annotate the type of the target object in the image to be annotated.

[0044] If the annotation labels assigned to the annotation file by the annotation model do not match the pre-configured annotation labels, it means that the annotation result is incorrect. In this case, the interactive information, the image to be annotated, and the incorrect annotation result are re-inputted into the annotation model. The annotation model will then re-annotate based on the input information to obtain a new annotation file. The annotation model will then continue to match the annotation labels assigned to the annotation file by the model with the preset annotation labels until the annotation labels assigned to the annotation file by the model match the preset annotation labels.

[0045] By verifying whether the annotation labels assigned to the annotation file by the annotation model match the preset annotation labels, the correct annotation results can be quickly filtered out, and the incorrect annotation results can be promptly input into the annotation model for re-annotation, which effectively improves the accuracy of annotation and thus improves the efficiency of annotation.

[0046] In one possible implementation, if the first annotation label matches at least one second annotation label, the output annotation file is then corrected until the annotation error rate, annotation category coverage, and pixel deviation of the annotation file meet preset conditions, and then the corrected annotation file is stored in the database.

[0047] Before annotation, the required annotation error rate, annotation category coverage, and pixel deviation can be pre-defined. The output annotation file is then corrected based on these pre-defined parameters. Correction methods include manual correction and automatic system correction. If the annotation task is simple and the target object is singular, automatic system correction can be used. If the annotated image is very complex and the system cannot perform automatic correction, manual correction is used. After correction, it is checked whether the corrected file meets the pre-defined annotation error rate, annotation category coverage, and pixel deviation. If it does, the annotation file is output and stored in a cloud database. If it does not meet the requirements, the unqualified annotation file continues to be corrected until the pre-defined annotation error rate, annotation category coverage, and pixel deviation are met.

[0048] Correcting the annotation results based on the pre-defined annotation error rate, annotation category coverage, and pixel deviation can yield more accurate annotation results. Inaccurate annotation results can be corrected again, avoiding inaccurate annotation results due to annotation errors in the annotation model, thus improving the accuracy of annotation and consequently improving annotation efficiency.

[0049] To better understand the image annotation method disclosed in this application, the following uses street view data as an example to provide a detailed explanation of the image annotation method provided in the embodiments of this application.

[0050] In related technologies, street view data is mainly labeled manually. For example, cars, roads, people, buildings and vegetation in the image are labeled manually. It can be seen that manual labeling of road data takes a long time, resulting in low efficiency of road data labeling.

[0051] To address the aforementioned problems, this application provides an image annotation method. The solution provided in this application can be used for the annotation of the aforementioned street view data. It is understood that the solution provided in this application can also be used for annotating images of people, special objects in architectural images, etc., which will not be elaborated upon here.

[0052] like Figure 2 As shown, Figure 2 This is a flowchart of an image annotation method provided in Embodiment 1 of this application, which includes steps 201-204:

[0053] Step 201: Obtain the image to be labeled.

[0054] The images to be labeled are street view photos or street view video frames.

[0055] Step 202: Process the image to be labeled using a pre-determined interaction strategy to obtain interaction information.

[0056] After acquiring the image to be labeled, it is processed through a pre-determined interaction strategy. Once the processing is complete, interaction information can be obtained, including the position of the labeled object, its edges, and the positive and negative points of the labeling results, etc.

[0057] In one possible implementation, before processing the image to be labeled using a pre-determined interaction strategy, the interaction strategy and labeling model can be determined based on the image labeling task, wherein different image labeling tasks correspond to different combinations of interaction strategies and labeling models.

[0058] Select the appropriate interaction strategy and annotation model combination based on the street view annotation task.

[0059] In another possible implementation, based on the image annotation task, at least one target object to be annotated is determined, and based on the type of the target object, an interaction strategy matching the type of the target object is determined from at least two preset interaction strategies.

[0060] Based on the street view annotation task, the target objects to be annotated are determined, and an interaction strategy matching the type of the target object is selected. Different interaction strategies are suitable for identifying the regions where different types of target objects are located in the image to be identified.

[0061] Step 203: Input the interaction information and the image to be labeled into the pre-trained labeling model to obtain the labeling file output by the labeling model.

[0062] The interactive information and the image to be labeled are input into a pre-trained labeling model. The labeling model can label the type of the target object in the input image and generate a label file to indicate the type of the target object.

[0063] In one possible implementation, such as Figure 3 As shown, Figure 3 This is a flowchart for generating the labeled model, including:

[0064] 301. Obtain the street view dataset.

[0065] Street view datasets can be public datasets or private datasets collected by data collection vehicles.

[0066] 302. Use the street view dataset to train the pre-built model to obtain the labeled model.

[0067] An annotation model is obtained through training, and the annotation model can generate annotation files to indicate the type of target objects.

[0068] 303. Deploy the labeled model to the cloud.

[0069] Deploying labeled models to the cloud eliminates limitations on location and device, supports collaborative work among multiple users, and provides overall control over the process from raw data to labeled results, thereby improving labeling efficiency.

[0070] 304. Encapsulate the annotation model deployed in the cloud.

[0071] It is encapsulated into an interface that can be input and output on the web to prepare for subsequent policy interaction and model reasoning.

[0072] Step 204: Check whether the first annotation label in the annotation file that indicates the type of the target object matches at least one preset second annotation label; if the first annotation label matches at least one second annotation label, output the annotation file.

[0073] Verify whether the annotation file correctly labels the type of the target object and whether it matches the preset annotation labels. If they match, the annotation result is correct, and the annotation file is output.

[0074] In one possible implementation, if the first annotation label does not match at least one second annotation label, the image to be annotated, the interaction information, and the annotation file are input into the annotation model to re-annotate the type of the target object in the image to be annotated.

[0075] If the first annotation label does not match at least one second annotation label, it indicates that the annotation result is incorrect. The interaction information, the image to be annotated, and the incorrect annotation result are then input into the annotation model. The annotation model re-annotates based on the input information to obtain a new annotation file. The verification process continues until the first annotation label matches at least one second annotation label.

[0076] Step 205: Correct the output annotation file until it meets the preset annotation error rate, annotation category coverage and pixel deviation, and store the corrected annotation file in the database.

[0077] The system presets the annotation error rate, annotation category coverage, and pixel deviation. If the corrected annotation file meets the preset annotation error rate, annotation category coverage, and pixel deviation, the corrected annotation file will be output and stored in the cloud database. If it does not meet the preset annotation criteria, the system will continue to correct the annotation until the preset annotation error rate, annotation category coverage, and pixel deviation are met.

[0078] like Figure 4 As shown, Figure 4 This is a schematic diagram of an image annotation device provided in an embodiment of this application. The device includes:

[0079] Module 401 is used to acquire the image to be labeled;

[0080] Image processing module 402 is used to process the image to be labeled through a predetermined interaction strategy to obtain interaction information, wherein the interaction information is used to indicate the position of the target object in the image to be labeled.

[0081] The annotation module 403 is used to input interactive information and the image to be annotated into a pre-trained annotation model to obtain an annotation file output by the annotation model, wherein the annotation file output by the annotation model is used to indicate the type of target object in the image to be annotated.

[0082] In one possible implementation, the image processing module 402 can also acquire image annotation tasks, and then determine the interaction strategy and annotation model based on the image annotation tasks, wherein different image annotation tasks correspond to different combinations of interaction strategies and annotation models.

[0083] In one possible implementation, the image processing module 402 can determine at least one target object to be labeled according to the image annotation task, and then determine an interaction strategy that matches the type of the target object from at least two preset interaction strategies according to the type of the target object.

[0084] In one possible implementation, the image annotation device may further include: an encapsulation module;

[0085] The encapsulation module can acquire a street view dataset including street view data, use the street view dataset to train a pre-built model to obtain an labeled model, deploy the labeled model to the cloud, and encapsulate the labeled model deployed in the cloud to obtain a web interface for calling the labeled model.

[0086] In one possible implementation, the image annotation device further includes: a verification module.

[0087] The verification module can verify whether the first annotation label in the annotation file, which indicates the type of the target object, matches at least one preset second annotation label; if the first annotation label matches at least one second annotation label, the annotation file is output.

[0088] In one possible implementation, if the first annotation label does not match at least one second annotation label, the verification module can input the image to be annotated, the interaction information, and the annotation file into the annotation model to re-annotate the type of the target object in the image to be annotated.

[0089] In one possible implementation, the image annotation device may further include: a correction module;

[0090] The correction module is used to correct the output annotation file when the first annotation label matches at least one second annotation label, until the annotation error rate, annotation category coverage and pixel deviation of the annotation file meet the preset conditions, and then the corrected annotation file is stored in the database.

[0091] It should be noted that the information interaction and execution process between the modules in the above-mentioned image annotation device are based on the same concept as the aforementioned simulation model calibration method embodiment. For details, please refer to the description in the aforementioned image annotation method embodiment, and it will not be repeated here.

[0092] Reference Figure 5 This document illustrates a schematic diagram of an electronic device according to an embodiment of this application. The specific embodiments of this application do not limit the specific implementation of the electronic device.

[0093] like Figure 5 As shown, the electronic device may include: a processor 502, a communications interface 504, a memory 506, and a communications bus 508.

[0094] in:

[0095] The processor 502, communication interface 504, and memory 506 communicate with each other via communication bus 508.

[0096] Communication interface 504 is used to communicate with other electronic devices or servers.

[0097] The processor 502 is used to execute program 510, which can specifically execute the relevant steps in the above-described image annotation method embodiment.

[0098] Specifically, program 510 may include program code that includes computer operation instructions.

[0099] The processor 502 may be a central processing unit (CPU), a graphics processing unit (GPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of this application. The one or more processors included in the smart device may be processors of the same type, such as one or more CPUs; one or more GPUs; or they may be processors of different types, such as one or more CPUs, one or more GPUs, and one or more ASICs.

[0100] Memory 506 is used to store program 510. Memory 506 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk storage device.

[0101] Specifically, program 510 can be used to cause processor 502 to execute the image annotation method in any of the foregoing embodiments.

[0102] The specific implementation of each step in procedure 510 can be found in the corresponding steps and units described in any of the foregoing image annotation method embodiments, and will not be repeated here. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the devices and modules described above can be referred to the corresponding process descriptions in the foregoing method embodiments, and will not be repeated here.

[0103] The electronic device in this embodiment processes the image to be labeled using a pre-determined interaction strategy to obtain interaction information. This interaction information, along with the image to be labeled, is then input into a pre-trained labeling model. Through inference by the labeling model, a labeled file indicating the type of target object in the image can be obtained. As can be seen, interaction information can be obtained through the interaction strategy, indicating the position of objects in the image. The labeling model quickly identifies the position of target objects in the image based on this interaction information, thereby identifying the type of target objects and generating a labeled file indicating the type of target objects. The labeling process requires no manual intervention, solving the problems of time-consuming and inefficient manual labeling, and achieving efficient labeling of massive amounts of road data.

[0104] This application also provides a computer program product, including computer instructions that instruct a computing device to perform an operation corresponding to any of the methods in the above-described multiple method embodiments.

[0105] It should be noted that, depending on the implementation needs, the various components / steps described in the embodiments of this application can be broken down into more components / steps, or two or more components / steps or parts of the operation of components / steps can be combined into new components / steps to achieve the purpose of the embodiments of this application.

[0106] The methods described in the embodiments of this application can be implemented in hardware, firmware, or as software or computer code that can be stored in a recording medium (such as a CD-ROM, RAM, floppy disk, hard disk, or magneto-optical disk), or as computer code downloaded over a network that is originally stored in a remote recording medium or a non-transitory machine-readable medium and will be stored in a local recording medium. Thus, the methods described herein can be processed by software stored on a recording medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware (such as an ASIC or FPGA). It is understood that the computer, processor, microprocessor controller, or programmable hardware includes storage components (e.g., RAM, ROM, flash memory, etc.) capable of storing or receiving software or computer code that, when accessed and executed by the computer, processor, or hardware, implements the image annotation methods described herein. Furthermore, when a general-purpose computer accesses code used to implement the image annotation methods shown herein, the execution of the code transforms the general-purpose computer into a dedicated computer for executing the image annotation methods shown herein.

[0107] Those skilled in the art will recognize that the units and method steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the embodiments of this application.

[0108] The above embodiments are only used to illustrate the embodiments of this application, and are not intended to limit the embodiments of this application. Those skilled in the art can make various changes and modifications without departing from the spirit and scope of the embodiments of this application. Therefore, all equivalent technical solutions also fall within the scope of the embodiments of this application, and the patent protection scope of the embodiments of this application should be defined by the claims.

Claims

1. An image annotation method, characterized in that, include: Obtain the image to be labeled; Obtain image annotation tasks; Based on the image annotation task, at least one target object to be annotated is determined; Based on the type of the target object, an interaction strategy matching the type of the target object is determined from at least two preset interaction strategies; the image to be labeled is processed using the determined interaction strategy to obtain interaction information, wherein the interaction information is used to indicate the position of the target object in the image to be labeled; the interaction information includes the position, edges, and positive and negative points of the labeling result of the labeled object; The interaction information and the image to be labeled are input into a pre-trained labeling model to obtain a labeling file output by the labeling model, wherein the labeling file output by the labeling model is used to indicate the type of target object in the image to be labeled; The method further includes: checking whether a first annotation label in the annotation file indicating the type of the target object matches at least one preset second annotation label; if the first annotation label matches the at least one second annotation label, then outputting the annotation file; The first label is the label assigned to the labeling file by the labeling model, and the second label is a pre-set label. Different labeling tasks correspond to different pre-set labels.

2. The method according to claim 1, characterized in that, The method further includes: Based on the image annotation task, the annotation model is determined, wherein different image annotation tasks correspond to different combinations of interaction strategies and annotation models.

3. The method according to claim 1, characterized in that, The method further includes: Obtain a street view dataset that includes street view data; The pre-built model is trained using the street view dataset to obtain the labeled model; The annotation model is deployed to the cloud, and the annotation model deployed in the cloud is encapsulated to obtain a web interface for calling the annotation model.

4. The method according to claim 1, characterized in that, The method further includes: If the first annotation label does not match the at least one second annotation label, the image to be annotated, the interaction information, and the annotation file are input into the annotation model to re-annotate the type of the target object in the image to be annotated.

5. The method according to claim 1, characterized in that, After outputting the annotation file, the method further includes: The output annotation file is corrected until the annotation error rate, annotation category coverage, and pixel deviation of the annotation file meet the preset conditions, and the corrected annotation file is stored in the database.

6. An image annotation device, characterized in that, The device includes: An acquisition module is used to acquire an image to be labeled; and acquire an image labeling task; determine at least one target object to be labeled according to the image labeling task; and determine an interaction strategy that matches the type of the target object from at least two preset interaction strategies according to the type of the target object. An image processing module is used to acquire an image annotation task; determine at least one target object to be annotated according to the image annotation task; determine an interaction strategy that matches the type of the target object from at least two preset interaction strategies according to the type of the target object; process the image to be annotated through the determined interaction strategy to obtain interaction information, wherein the interaction information is used to indicate the position of the target object in the image to be annotated; the interaction information includes the position, edges, and positive and negative points of the annotation result of the annotated object; The annotation module is used to input the interaction information and the image to be annotated into a pre-trained annotation model to obtain an annotation file output by the annotation model, wherein the annotation file output by the annotation model is used to indicate the type of target object in the image to be annotated; The image annotation device further includes: a verification module: the verification module is used to verify whether a first annotation label indicating the type of the target object in the annotation file matches at least one preset second annotation label; if the first annotation label matches at least one second annotation label, the annotation file is output; The first label is the label assigned to the labeling file by the labeling model, and the second label is a pre-set label. Different labeling tasks correspond to different pre-set labels.

7. An electronic device, characterized in that, include: The processor, memory, communication interface, and communication bus are provided, wherein the processor, memory, and communication interface communicate with each other via the communication bus. The memory is used to store at least one executable instruction that causes the processor to perform the operation corresponding to the method as described in any one of claims 1-5.

8. A computer storage medium, characterized in that, It stores a computer program that, when executed by a processor, implements the method as described in any one of claims 1-5.