Image labeling method and device, model training method and device and electronic equipment
By providing an image annotation method, using the object detection model to automatically mark and intelligently add labels, the problem of low image annotation efficiency in the prior art is solved, and efficient automation of image annotation is achieved.
Patent Information
- Application Number
- CN202510590431.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-08
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-05-08
AI Technical Summary
In the prior art, image labeling efficiency is low and requires a lot of manual time to complete. Multiple images are marked and multiple target objects in each image are large in size and low in efficiency.
An image annotation method is provided, by displaying the training image to be marked, receiving user annotations, using the marked images to train the object detection model, automatically mark and intelligently add labels, and reducing the amount of manual annotation.
It improves the efficiency of image labeling and reduces the cost of manual labeling. Through automatic marking and intelligent labeling functions, efficient automation of image labeling is achieved.
Smart Images

Figure CN120126141A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of machine vision technology, and particularly to an image annotation method, a model training method, an apparatus, and an electronic device. Background Art
[0002] For a target detection model, image annotation refers to bounding the positions of target objects of interest in an image and adding class labels indicating their categories, so that when training the target detection model, the model can accurately extract the features of the target objects.
[0003] However, currently, the work of annotating images usually needs to be done manually. In a single training process of a target detection model, it is usually necessary to annotate multiple images, and there may be multiple target objects in each image. Therefore, the amount of annotation required is large, resulting in low annotation efficiency and consuming a large amount of time cost. Summary of the Invention
[0004] The purpose of the embodiments of this application is to provide an image annotation method, a model training method, an apparatus, and an electronic device to improve the efficiency of image annotation. The specific technical solutions are as follows:
[0005] In a first aspect, the embodiments of this application provide an image annotation method applied to an electronic device. The method includes:
[0006] Display a first training image to be annotated;
[0007] Receive the annotation of the user for the first training image, and display the target boxes and class labels of different target objects annotated by the user on the first training image; wherein, each class label annotates at least one target object;
[0008] In response to an automatic labeling instruction, display on a second training image: the target boxes and class labels of the target objects in the second training image; wherein, the target boxes and class labels of the target objects in the second training image are obtained by inputting the second training image into a target detection model trained with the annotated first training image; the annotated first training image includes: the target boxes and class labels annotated by the user.
[0009] In one embodiment, after receiving the annotation of the user for the first training image and displaying the target boxes and class labels of different target objects annotated by the user on the first training image, the method further includes:
[0010] In response to a label append instruction, display on the first training image: the target boxes and class labels of the target objects in the first training image;
[0011] Among them, the target bounding box and class label of the target object in the first training image are obtained by inputting the first training image into a target detection model trained with the first training image annotated by the user;
[0012] The annotated first training image further includes: the target bounding box and class label annotated after additional annotation.
[0013] In one embodiment, before receiving the annotation of the user for the first training image, the method further includes:
[0014] In response to a label creation instruction, receive and save different class labels corresponding to different target objects created by the user for the first training image.
[0015] In one embodiment, when the number of second training images to be annotated is greater than 1, after executing the automatic labeling instruction on the first second training image, the method further includes:
[0016] Sequentially display the next second training image to be annotated, and execute the automatic labeling instruction for this second training image until all second training images are labeled.
[0017] In one embodiment, before the step of displaying the first training image to be annotated, the method further includes:
[0018] Obtain a training image set determined by the user; wherein, the training image set includes multiple training images;
[0019] Receive a training image selected by the user from the training image set as the first training image to be annotated; wherein, except for this first training image to be annotated in the training image set, other images are used as the second training images to be annotated.
[0020] In one embodiment, for the first second training image, in response to the automatic labeling instruction, display on the second training image: the target bounding box and class label of the target object in this second training image, including:
[0021] In response to the automatic labeling instruction, display on the second training image the target bounding box and class label of the target object annotated after the second training image is input into the target detection model; and,
[0022] Based on the user's selection, display on the second training image the target bounding box and class label of all target objects after the user's correction and adjustment.
[0023] In one embodiment, after each execution of the automatic labeling instruction on a second training image, the method further includes: based on the user's selection, add the second training image corrected and adjusted by the user to the target training set to update the target training set;
[0024] After each display of the next second training image to be labeled and before responding to the automatic labeling instruction, the method further includes: training the object detection model with the updated target training set based on user selection to update the object detection model;
[0025] For the second and subsequent second training images, in response to the automatic labeling instruction, display on the second training image: the target bounding boxes and class labels of the target objects in the second training image, including:
[0026] Based on user selection, display on the second training image the target bounding boxes and class labels of the target objects labeled after the second training image is input into the unupdated or updated object detection model; and,
[0027] Based on user selection, display on the second training image the target bounding boxes and class labels of all target objects after user correction and adjustment.
[0028] In one embodiment, in response to the label append instruction, display on the first training image: the target bounding boxes and class labels of the target objects in the first training image, including:
[0029] In response to the label append instruction, display on the first training image the target bounding boxes and class labels of the target objects newly labeled this time; wherein, the target bounding boxes and class labels of the target objects newly labeled this time are obtained by inputting the first training image into the object detection model trained with the first training image labeled by the user, and the similarity between the target objects newly labeled this time and the target objects labeled by the user is greater than a preset threshold.
[0030] In one embodiment, after the label append instruction is executed and when not all the target bounding boxes and class labels of the target objects are displayed on the first training image, the method further includes:
[0031] In response to the next label append instruction, display on the first training image the target bounding boxes and class labels of the target objects newly labeled this time; wherein, the target bounding boxes and class labels of the target objects newly labeled this time are obtained by inputting the first training image into the updated object detection model, and the similarity between the target objects newly labeled this time and the already labeled target objects is greater than a preset threshold;
[0032] The updated object detection model is obtained by training the object detection model based on the first training image including the target bounding boxes and class labels labeled by the user and the target bounding boxes and class labels newly labeled before.
[0033] In one embodiment, after displaying the target box and class label of the target object for the current additional annotation, the method further includes:
[0034] Based on the user's selection, correcting and adjusting the target box and class label of the target object for the current additional annotation.
[0035] In one embodiment, before responding to the label creation instruction, the method further includes: displaying a label creation control; wherein, the label creation instruction is generated by the user clicking on the label creation control;
[0036] Before receiving the annotation of the user for the first training image, the method further includes: displaying label controls corresponding to different target objects created by the user; wherein, the target box and class label of each target object annotated by the user on the first training image are obtained by annotating under the condition that the user selects the label control corresponding to the target object;
[0037] Before responding to the automatic labeling instruction, the method further includes: displaying an automatic labeling control; wherein, the automatic labeling instruction is generated by the user clicking on the automatic labeling control;
[0038] Before responding to the label addition instruction, the method further includes: displaying an intelligent addition control; wherein, the label addition instruction is generated by the user clicking on the intelligent addition control.
[0039] In one embodiment, the responding to the automatic labeling instruction includes: responding to the automatic labeling instruction when the automatic labeling control is unlocked; wherein, the unlocking condition of the automatic labeling control is: the first training image has been annotated and the second training image to be annotated is displayed;
[0040] The responding to the label addition instruction includes: responding to the label addition instruction when the intelligent addition control is unlocked; wherein, the unlocking condition of the intelligent addition control is: there are target boxes and class labels annotated by the user on the first training image.
[0041] In one embodiment, the method further includes:
[0042] When a modification instruction for any target box is received, modifying the position, size, and / or corresponding class label of the target box based on the modification instruction;
[0043] When a deletion instruction for any target box is received, removing the target box and the corresponding class label from the training image;
[0044] When new annotations added by the user for the training image are received, displaying the new target boxes and class labels added by the user on the training image.
[0045] In a second aspect, an embodiment of the present application provides an image annotation method applied to an electronic device. The method includes:
[0046] Display a training image to be annotated.
[0047] Receive an annotation from the user for the training image, and display bounding boxes and class labels of different target objects annotated by the user on the training image; wherein, each class label annotates at least one target object.
[0048] In response to a label append instruction, display on the training image: the bounding boxes and class labels of the target objects in the training image.
[0049] Among them, the bounding boxes and class labels of the target objects in the training image are obtained by inputting the training image into a target detection model trained with the training images annotated by the user.
[0050] In one embodiment, the number of training images to be annotated is multiple; after executing the label append instruction for the first training image, the method further includes:
[0051] Sequentially display the next training image to be annotated, and execute the label append instruction for this training image until all training images are labeled.
[0052] In one embodiment, after executing the label append instruction for the first training image, the method further includes:
[0053] Sequentially display the next training image to be annotated.
[0054] In response to an automatic labeling instruction, display on the next training image: the bounding boxes and class labels of the target objects in this training image; wherein, the bounding boxes and class labels of the target objects in this training image are obtained by inputting this training image into a target detection model trained with the already annotated training images.
[0055] Return to execute the step of sequentially displaying the next training image to be annotated until all training images are labeled.
[0056] In a third aspect, an embodiment of the present application provides a model training method applied to an electronic device. The method includes:
[0057] Obtain the training images determined by the user.
[0058] Apply the image annotation method described in any of the above to annotate the training images.
[0059] In response to a model training instruction, a target detection model is trained using the labeled training images to obtain a trained target detection model.
[0060] In a fourth aspect, an embodiment of the present application provides an image annotation device applied to an electronic device. The device includes:
[0061] A display module for displaying a first training image to be annotated;
[0062] An annotation receiving module for receiving an annotation of the user for the first training image and displaying bounding boxes and class labels of different target objects annotated by the user on the first training image; wherein, each class label annotates at least one target object;
[0063] An automatic labeling response module for, in response to an automatic labeling instruction, displaying on a second training image: bounding boxes and class labels of target objects in the second training image; wherein, the bounding boxes and class labels of target objects in the second training image are obtained by inputting the second training image into a target detection model trained using the labeled first training image; the labeled first training image includes: bounding boxes and class labels annotated by the user.
[0064] In a fifth aspect, an embodiment of the present application provides an image annotation device applied to an electronic device. The device includes:
[0065] A training image display module for displaying a training image to be annotated;
[0066] A user annotation receiving module for receiving an annotation of the user for the training image and displaying bounding boxes and class labels of different target objects annotated by the user on the training image; wherein, each class label annotates at least one target object;
[0067] A label append instruction response module for displaying on the training image: bounding boxes and class labels of target objects in the training image; wherein, the bounding boxes and class labels of target objects in the training image are obtained by inputting the training image into a target detection model trained using the training image annotated by the user.
[0068] In a sixth aspect, an embodiment of the present application provides a model training device applied to an electronic device. The device includes:
[0069] A training image obtaining module for obtaining a training image determined by the user;
[0070] A training image annotation module for annotating the training image by applying any one of the above-mentioned image annotation methods;
[0071] A model training instruction response module, which is configured to respond to a model training instruction, and use the labeled training images to train a target detection model to obtain a trained target detection model.
[0072] In a seventh aspect, an embodiment of the present application provides an electronic device, including:
[0073] A memory for storing a computer program;
[0074] A processor, when executing the program stored in the memory, implements any one of the above-described image annotation methods or model training methods.
[0075] In an eighth aspect, an embodiment of the present application provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, it implements any one of the above-described image annotation methods or model training methods.
[0076] In a ninth aspect, an embodiment of the present application further provides a computer program product containing instructions, which when running on a computer, causes the computer to execute any one of the above-described image annotation methods or model training methods.
[0077] Advantages of the embodiments of the present application:
[0078] Applying the embodiments of the present invention, a user can first manually annotate the first training image displayed on the electronic device, so that the electronic device can use the labeled first training image to train a target detection model; in this way, the user can send an automatic labeling instruction to make the electronic device input the second training image into the trained target detection model, so as to automatically obtain the target box and class label of the target object in the second training image. It can be seen that applying the embodiments of the present invention only requires manual annotation of the first training image, and the second training image can be automatically annotated by the electronic device. Therefore, through this solution, images can be annotated more efficiently, reducing the time cost of annotation.
[0079] Of course, implementing any product or method of the present application does not necessarily require achieving all the above advantages at the same time. Description of the Drawings
[0080] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application, and those of ordinary skill in the art can also obtain other embodiments based on these drawings.
[0081] Figure 1 It is a structural schematic diagram of a few-shot target detection training operator in the embodiments of the present application;
[0082] Figure 2 Schematic diagram of the principle process for training an object detection model using the few-shot object detection training operator shown in Figure 1 ;
[0083] Figure 3 Schematic diagram of the process of the first embodiment of the first annotation method for images provided by the embodiments of the present application;
[0084] Figure 4 Schematic diagram of the process of the second embodiment of the first annotation method for images provided by the embodiments of the present application;
[0085] Figure 5(a) is a first schematic diagram of the interactive interface used in the annotation method for images provided by the embodiments of the present application (state without importing an image);
[0086] Figure 5(b) is a second schematic diagram of the interactive interface shown in Figure 5(a) (state of displaying common parameters for object detection);
[0087] Figure 6 is a third schematic diagram of the interactive interface shown in Figure 5(a) (state of importing an image);
[0088] Figure 7 is a fourth schematic diagram of the interactive interface shown in Figure 5(a) (state when creating a label);
[0089] Figure 8 is a fifth schematic diagram of the interactive interface shown in Figure 5(a) (state after label creation is completed);
[0090] Figure 9 is a sixth schematic diagram of the interactive interface shown in Figure 5(a) (state after the user annotates the first training image);
[0091] Figure 10 is a seventh schematic diagram of the interactive interface shown in Figure 5(a) (state of displaying the automatic labeling button);
[0092] Figure 11 is Figure 4 Schematic diagram of the process of intelligent addition in the embodiment shown in
[0093] Figure 12(a) is a schematic diagram of label addition in the annotation method for images provided by the embodiments of the present application;
[0094] Figure 12(b) is a schematic diagram in which all target objects are marked in the annotation method for images provided by the embodiments of the present application;
[0095] Figure 13 is Figure 4 Schematic diagram of the intelligent addition interaction principle of the embodiment shown in
[0096] Figure 14 is Figure 4 Schematic diagram of the automatic annotation process in the illustrated embodiment;
[0097] Figure 15 is the eighth schematic diagram of the interaction interface shown in Fig. 5(a) (status of prompting to reject an image from being added to the target training set);
[0098] Figure 16 is the ninth schematic diagram of the interaction interface shown in Fig. 5(a) (status of prompting whether to retrain the model);
[0099] Figure 17 is Figure 4 Schematic diagram of the automatic marking interaction principle of the illustrated embodiment;
[0100] Figure 18 is the schematic flowchart of the first embodiment of the second annotation method for images provided by the embodiments of the present application;
[0101] Figure 19 is the schematic flowchart of the second embodiment of the second annotation method for images provided by the embodiments of the present application;
[0102] Figure 20 is the schematic flowchart of the third embodiment of the second annotation method for images provided by the embodiments of the present application;
[0103] Figure 21 is the schematic flowchart of the embodiment of the model training method provided by the embodiments of the present application;
[0104] Figure 22 is the schematic structural diagram of the first embodiment of the image annotation device provided by the embodiments of the present application;
[0105] Figure 23 is the schematic structural diagram of the second embodiment of the image annotation device provided by the embodiments of the present application;
[0106] Figure 24 is the schematic structural diagram of the model training device provided by the embodiments of the present application;
[0107] Figure 25 is the schematic structural diagram of the electronic device provided by the embodiments of the present application. Detailed implementation manners
[0108] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art based on the present application belong to the scope of protection of the present application.
[0109] In this embodiment, the target detection model to be trained can be a deep learning model, such as a few-shot object detection model (Few-Shot Object Detection, FSOD). In order to train this model, a dedicated training program can be developed, and this program can be called a few-shot object detection training operator. Refer to Figure 1 , Figure 1 which is a schematic structural diagram of the few-shot object detection training operator in the embodiments of the present application; as Figure 1 shown, the following contents can be encapsulated in this program: a basic target detection deep learning model (that is, the target detection model to be trained, hereinafter referred to as the target detection model), common parameters for target detection, and a target detection annotation function. For example, the common parameters for target detection can include: angle enable, which is used for the user to determine whether to use a rectangular box with an angle value to annotate the target object; optimal model resolution setting enable, which is used for the user to determine whether to adjust the resolution of the displayed training image; and the optimal model resolution, which is used for the user to set the resolution of the displayed training image.
[0110] The user can configure the common parameters of the target detection model through this program, mark the target object of interest in the training image through the target detection annotation function, and finally train the basic target detection deep learning model according to the marked training image to obtain a target detection model suitable for the current actual scenario. Of course, in addition to the few-shot object detection model, the target detection model to be trained can also be other models, such as the YOLO (You Only Look Once) model, the SSD (Single Shot MultiBox Detector) model, etc.
[0111] For the above-mentioned target detection annotation function, a method needs to be provided to improve the efficiency of annotating the training image, and thus the efficiency of training the target detection model can be improved. Therefore, the embodiments of the present application provide an image annotation method, a model training method, a device, and an electronic device. First, the image annotation method provided by the embodiments of the present application will be introduced below. This method can be applied to an electronic device, and the electronic device can be integrated with a display screen to display the interactive interface in this embodiment. For example, the electronic device can be a computer, a tablet computer, a smart phone, etc.
[0112] See Figure 2 , Figure 2 which is a schematic diagram of the principle process for training an object detection model using the few-shot object detection training operator shown Figure 1 ; as shown Figure 2 in
[0113] S201, Add 1 - 10 training images with significant feature differences;
[0114] S202, Create corresponding numbers of object type labels according to the requirements of the actual scenario;
[0115] S203, Create an object example for each type of object;
[0116] S204, Mark all objects in the remaining images through "intelligent append" or "automatic labeling";
[0117] S205, Train the images added to the object training set (based on a deep learning model).
[0118] During the training process of the deep learning model, in order to improve the annotation accuracy and cover sufficient actual scenarios, hundreds of example images are required as samples for training, which is time-consuming and laborious. Although ordinary few-shot learning training reduces the number of samples needed, the repetitive annotation operations are still inevitable. The embodiment of this application combines the annotation methods of few-shot detection training, intelligent append, and automatic annotation, and can annotate multiple images and multiple target objects, effectively improving the work efficiency of object detection annotation in complex scenarios.
[0119] Based on Figure 2 the principle shown
[0120] the embodiments of the present invention provide two specific annotation methods: the first is to annotate the training images by combining manual annotation with functions such as automatic labeling; the second is to annotate the training images by combining manual annotation with the intelligent append function.
[0121] First, a detailed description of the first method for annotating training images will be given.
[0122] See Figure 3 , Figure 3 which is a schematic diagram of the process of the first embodiment of the first annotation method for images provided by the embodiments of this application. As shown Figure 3 in
[0123] S301, Display the first training image to be annotated;
[0124] S302. Receive the annotation of the user for the first training image, and display the target boxes and class labels of different target objects annotated by the user on the first training image; wherein, each class label annotates at least one target object.
[0125] S303. In response to the automatic labeling instruction, display on the second training image: the target boxes and class labels of the target objects in the second training image; wherein, the target boxes and class labels of the target objects in the second training image are obtained by inputting the second training image into a target detection model trained with the annotated first training image; the annotated first training image includes: the target boxes and class labels annotated by the user.
[0126] In this embodiment, the user can first perform manual annotation on the first training image displayed by the electronic device, so that the electronic device can train the target detection model with the annotated first training image; in this way, the user can make the electronic device input the second training image into the trained target detection model through the automatic labeling instruction, so as to automatically obtain the target boxes and class labels of the target objects in the second training image. It can be seen that applying the embodiment of the present invention only requires manual annotation of the first training image, and the second training image can be automatically annotated by the electronic device. Therefore, through this solution, the image can be annotated more efficiently, and the time cost of annotation can be reduced.
[0127] In some embodiments, if the training set contains both the first training image and the second training image, before automatically labeling the second training image, the user can also perform intelligent addition to the first training image to display more target boxes and class labels on the first training image, and then automatically label the second training image.
[0128] Specifically, refer to Figure 4 , Figure 4 which is a schematic flowchart of the second embodiment of the first annotation method of the image provided by the embodiment of the present application. As Figure 4 shown, the method includes:
[0129] S401. Obtain the training set determined by the user; wherein, the training set includes multiple training images.
[0130] In this embodiment, the electronic device can display an interactive interface for annotating training images. Referring to FIG. 5(a), FIG. 5(a) is a first schematic diagram of the interactive interface used in the image annotation method provided by the embodiment of the present application (state without importing an image). As shown in FIG. 5(a), in the area of "1 / Add Training Image", the user can click "Camera Capture" to control the camera connected to the electronic device to capture training images in real time; or click "Stored Image Import" to select an image from the images pre-stored in the electronic device as a training image; or click "External Import" to obtain a training image from an external device connected to the electronic device.
[0131] While displaying the interactive interface shown in FIG. 5(a), the electronic device can also display a dialog box for setting common parameters for object detection. Referring to FIG. 5(b), FIG. 5(b) is a second schematic diagram of the interactive interface used in the image annotation method provided by the embodiment of the present application (state of displaying common parameters for object detection). As shown in FIG. 5(b), the dialog box includes: a switch control corresponding to angle enable. The user can click the switch control to turn on or off the angle enable. When the angle enable is turned on, in the training images of the object detection model used for training, rectangular boxes with angle values will be used to annotate the target objects; when it is turned off, rectangular boxes without angles will be used to annotate the target objects. It can be understood that the rectangular box without an angle refers to a horizontal and vertical rectangular box.
[0132] The dialog box may also include: a switch control corresponding to the enablement of the optimal model resolution setting. This parameter can be default not enabled. If the user believes that the display effect of the current training image is not good, the user can click the switch control to turn on this parameter. After turning it on, the parameter "optimal model resolution" will be displayed in the dialog box. Then, the user can set the value of the optimal model resolution in the dialog box. In this way, when the subsequent interactive interface displays the training image, the number of pixels on the longer side of the displayed training image will be adjusted according to the value set by the user, and the other side will be scaled proportionally. For example, if the set value is 512, the number of pixels on the longer side of the displayed training image is 512.
[0133] When the object detection model to be trained is a few-shot object detection model, if the first method for annotating training images is adopted, the user can select 2-10 training images as the training image set. Of course, it is also possible to have more than 10 training images. The user can try to select images with significant feature differences to improve the diversity of the samples, thereby improving the training effect of the object detection model. Each training image can contain only one type of target object or multiple types. If the second method for annotating training images is adopted, the user can select 1-10 training images as the training image set. The specific process will be described in detail below.
[0134] S402, receive a training image selected by the user from the training image set as the first training image to be labeled. Among them, except for the first training image to be labeled in the training image set, other images are used as the second training images to be labeled.
[0135] After the user determines the training image set, the user can select an image including all category target objects as the first training image. By obtaining the training image set determined by the user and receiving a training image selected by the user from the training image set as the first training image to be labeled, a basis can be provided for subsequent labeling of the first training image and the second training images.
[0136] S403, display the first training image to be labeled.
[0137] In this step, the user can select an image as the first training image after obtaining the training image set determined by the user. After the user selects the first training image, the first training image can be displayed in the user interaction interface of the electronic device. For other training images that are not currently displayed, the electronic device can display their thumbnails in the thumbnail display area. See Figure 6 , Figure 6 as the third schematic diagram (state of importing an image) of the interaction interface shown in Fig. 5(a). As Figure 6 shown, the thumbnail corresponding to the currently displayed training image can show the effect of a stroke to prompt the user which training image is currently displayed.
[0138] S404, in response to the label creation instruction, receive and save different category labels corresponding to different target objects created by the user for the first training image.
[0139] Before responding to the label creation instruction, the electronic device can display a label creation control. In this way, the label creation instruction can be generated by the user clicking the label creation control. Still taking Figure 6 as an example, the label creation control can be "New" in the "Label List" area. After that, the electronic device can display a label creation dialog box. See Figure 7 , Figure 7 as the fourth schematic diagram (state when creating a label) of the interaction interface shown in Fig. 5(a). As Figure 7 shown, the user can enter a label name and set a label color in this dialog box, so that different category labels can be displayed in different colors, which is convenient for the user to distinguish.
[0140] For example, if there are two different target objects in the first training image: large pills and small pills, the user can create two category labels and enter the label names as "large pills" and "small pills" respectively. After the user creates the category labels, the corresponding label controls for the created category labels can be displayed below "Label Category". Refer to Figure 8 , Figure 8 which is the fifth schematic diagram of the interactive interface shown in Fig. 5(a) (the state after label creation). As Figure 8 shown, the label control corresponding to each category label can also include "Edit" and "Delete" buttons, which are used to modify the name of the category label and delete the category label respectively.
[0141] In this embodiment, by receiving instructions through the display control, it can enable the user to trigger the corresponding functions more conveniently and improve the user experience.
[0142] S405, Receive the annotation of the first training image by the user, and display the target boxes and category labels of different target objects annotated by the user on the first training image; wherein, each category label annotates at least one target object;
[0143] Taking Figure 8 as an example, the user can click "Rectangular Selection" in the "2 / Annotated Image" area, and then annotate the target objects in the first training image. For example, the user can draw a rectangle at the position of the target object as the target box to enclose the target object inside the rectangle, and the electronic device will display the target box drawn by the user, and at the same time display the label category of the target object near the target box.
[0144] Before receiving the annotation of the first training image by the user, the electronic device can first display the label controls corresponding to different target objects created by the user. In this case, the target box and category label of each target object annotated by the user on the first training image are obtained by annotation when the user selects the label control corresponding to the target object.
[0145] When there are multiple target objects in the first training image, the user can first select the label control corresponding to one of the created target objects, and the selected label control can also be displayed in a stroked form. Then the user can annotate the target object in the training image. After that, the user can also select the label controls corresponding to other created target objects to annotate the target objects of other categories, so as to realize the annotation of different categories of target objects. Refer to Figure 9 , Figure 9 which is the sixth schematic diagram of the interactive interface shown in Fig. 5(a) (the state after the user annotates the first training image). As Figure 9As shown, the user can mark a "big pill" and a "small pill" in the image respectively. At this time, the target boxes and class labels marked by the user can be used as target examples. For each type of target object, the user only needs to mark one. Of course, the user can also mark multiple target objects for each category to effectively enrich the target examples and make subsequent additional annotations more accurate.
[0146] S406, in response to the label append instruction, display on the first training image: the target box and class label of the target object in the first training image;
[0147] Among them, the target box and class label of the target object in the first training image are obtained by inputting the first training image into the target detection model trained with the first training image marked by the user;
[0148] Before responding to the label append instruction, the electronic device can first display an intelligent append control; in this way, the label append instruction can be generated by the user clicking the intelligent append control. When the user clicks the intelligent append control, there is a certain time-consuming in responding to the label append instruction. During this period, the electronic device can also provide a certain loading waiting prompt, for example, it can display a dialog box containing text to prompt the user.
[0149] In one implementation, responding to the label append instruction includes: responding to the label append instruction when the intelligent append control is unlocked; among them, the unlocking condition of the intelligent append control is: there are target boxes and class labels marked by the user on the first training image.
[0150] That is to say, when there are no target boxes and class labels marked by the user in the training set, the intelligent append control can be in a locked state. For example, the electronic device can temporarily not display the intelligent append control, or display the intelligent append control in a lighter color until the intelligent append control is unlocked and then display the intelligent append control normally. By locking and unlocking the control, it is possible to prevent the user from accidentally touching the control when the function corresponding to the control cannot take effect.
[0151] If there are still unmarked target objects in the first training image marked by the user, the user can issue a label append instruction. After the electronic device receives the label append instruction, it can first train the target detection model based on the target boxes and class labels marked by the user, and then input the first training image into the target detection model to obtain and display the target boxes and class labels of the target objects detected by the target detection model in the first training image.
[0152] The target detection model can be trained by the method of few-shot detection training. Exemplarily, the training process of the target detection model can be as follows:
[0153] First, input the unannotated first training image into the target detection model to be trained, and obtain the bounding boxes and class labels of the target objects detected by the model for this image as the model prediction results. Calculate the model loss based on the model prediction results and the ground truth of the first training image, and use the calculated model loss to adjust the parameters of the target detection model until the model converges, obtaining the trained target detection model. Among them, the ground truth of the first training image can be the bounding boxes and class labels annotated in the first training image.
[0154] S407, display the second training image to be annotated;
[0155] After completing the annotation of the first training image, the user can switch to the second training image, so that the electronic device displays the second training image. For example, as Figure 6 shown, the user can click on the thumbnail of the currently unshown training image in the thumbnail display area to achieve the switching of the training images.
[0156] S408, in response to the automatic labeling instruction, display on the second training image: the bounding boxes and class labels of the target objects in the second training image;
[0157] Among them, the bounding boxes and class labels of the target objects in the second training image are obtained by inputting the second training image into the target detection model trained with the annotated first training image. The annotated first training image includes: the bounding boxes and class labels annotated by the user. Of course, in the case where the electronic device responds to the label addition instruction for the first training image, the annotated first training image also includes: the bounding boxes and class labels annotated after addition.
[0158] Similarly, before responding to the automatic labeling instruction, the electronic device can first display the automatic labeling control, so that the automatic labeling instruction can be generated by the user clicking on the automatic labeling control.
[0159] The step of responding to the automatic labeling instruction may include: responding to the automatic labeling instruction when the automatic labeling control is unlocked. Among them, the unlocking condition of the automatic labeling control is: the first training image has been annotated and the second training image to be annotated is displayed.
[0160] In one implementation, refer to Figure 10 , the seventh schematic diagram of the interactive interface shown in Fig. 5(a) (displaying the state of the automatic labeling button), as Figure 10As shown, while displaying the second training image, a Toast (a non-modal pop-up window) prompt box can be generated, and at the same time, an automatic labeling control (i.e., "Train and Automatically Label" in the figure) is displayed in the prompt box, and the text "The figure can be automatically labeled through automatic reasoning" is displayed. Of course, the specific presentation method of the automatic labeling control can also be defined according to the business scenario by itself, and this is only an example here.
[0161] After the electronic device receives the automatic labeling instruction, it can first use the labeled first training image to train the object detection model, and then input the second training image into the trained object detection model to obtain the target box and class label of the target object in the second training image and display them. Since this process takes a certain amount of time, the electronic device can also provide a certain loading waiting prompt during this period.
[0162] In a specific implementation manner, step S408 may include:
[0163] In response to the automatic labeling instruction, on the second training image, display the target box and class label of the target object labeled after the second training image is input into the object detection model; and, based on user selection, on the second training image, display the target box and class label of all target objects after the user's correction and adjustment.
[0164] Specifically, the process of correction and adjustment may include the following situations:
[0165] Situation 1, when a modification instruction for any target box is received, modify the position, size, and / or corresponding class label of the target box based on the modification instruction;
[0166] Situation 2, when a deletion instruction for any target box is received, remove the target box and the corresponding class label from the training image;
[0167] Situation 3, when a new annotation added by the user for the training image is received, display the new target box and class label added by the user on the training image.
[0168] In this embodiment, the electronic device supports the user to correct and adjust the target box and class label of the target object, which can further ensure the accuracy of the annotation, and thus can improve the training effect of the object detection model.
[0169] In the case where the number of second training images to be labeled is greater than 1, after the automatic labeling instruction is executed for the first second training image, the method may further include:
[0170] S409, sequentially display the next second training image to be labeled, and execute the automatic labeling instruction for the second training image until all second training images are labeled.
[0171] That is to say, the above steps S407 - S408 can be executed multiple times until all the second training images are labeled. In this way, in the case of multiple training images, each training image can also be quickly labeled. It can be seen that this solution provides an automatic labeling function applicable to multi - image scenarios, which can quickly identify the same type of existing targets under different images, effectively saving the labeling time for a large number of repeated targets in complex scenarios and improving the labeling efficiency.
[0172] Moreover, in this embodiment, by combining the few - shot detection training, intelligent addition, and automatic annotation of deep - learning samples in advance, the order of magnitude of the number of manual annotations is reduced from several hundreds or thousands to several or dozens, greatly reducing the time and effort consumed by manual annotation during target detection training, and also significantly reducing the time and labor costs for actual application deployment.
[0173] In this embodiment, the user can create labels in the electronic device and then manually annotate the first training image displayed on the electronic device, enabling the electronic device to train the target detection model using the labeled first training image. In this way, the electronic device can automatically annotate the target objects that the user has not annotated in the first training image. After the second training image is displayed, the user can use the automatic labeling instruction to make the electronic device input the second training image into the trained target detection model, thereby automatically obtaining the target boxes and class labels of the target objects in the second training image. It can be seen that through this solution, only some of the target objects in the first training image need to be manually annotated, and the remaining target objects in the first training image and the target objects in the second training image can be automatically annotated by the electronic device. Therefore, through this solution, images can be labeled more efficiently, reducing the time cost of labeling.
[0174] In an embodiment of the present application, refer to Figure 11 , Figure 11 For Figure 4 the schematic diagram of the intelligent addition process in the shown embodiment, as Figure 11 shown, this process may include the following steps:
[0175] S1101, in response to the label addition instruction, display the target box and class label of the target object to be added and labeled this time on the first training image; wherein, the target box and class label of the target object to be added and labeled this time are obtained by inputting the first training image into the target detection model trained using the first training image labeled by the user, and the similarity between the target object to be added and labeled this time and the target object labeled by the user is greater than a preset threshold.
[0176] Among them, the similarity between two target objects may refer to the image similarity between the partial images framed by the target boxes annotating the two target objects. That is to say, after the electronic device responds to the label append instruction, it may not append the target boxes and class labels of all the target objects detected by the target detection model, but only append the annotation of the target objects whose similarity to the target objects annotated by the user is greater than a preset threshold. Specifically, the electronic device may calculate the similarity between each target object under each category detected by the target detection model and the target objects of this category that have been annotated currently. If there are multiple target objects of this category that have been annotated currently, the average value of the similarities may be taken. In the case where the similarity is greater than the preset threshold, the target box and class label of this target object are appended for annotation, otherwise, no annotation is appended.
[0177] The preset threshold can be set according to the actual situation and experience. In one implementation, the preset threshold can be set relatively high, such as 98%, so that only 1-3 target objects are automatically appended for annotation each time. This quantity is convenient for the user to confirm and adjust, gradually releasing a small number of inference targets, which helps to continuously improve the inference accuracy.
[0178] Exemplarily, referring to FIG. 12(a), FIG. 12(a) is a schematic diagram of label append in the image annotation method provided by an embodiment of the present application. After the first append annotation is performed on the Figure 9 first training image therein, the obtained first training image can be as shown in FIG. 12(a).
[0179] After each append annotation, the electronic device can also, based on the user's selection, correct and adjust the target box and class label of the target object appended for annotation this time.
[0180] After the label append instruction is executed, step S1102 is performed: determining whether the target boxes and class labels of all target objects are displayed on the first training image; if not, S1103 is executed; if so, S1104 is executed.
[0181] In the case where the target boxes and class labels of all target objects are not displayed on the first training image, the method may further include:
[0182] S1103, in response to the next label append instruction, on the first training image, displaying the target boxes and class labels of the target objects appended for annotation this time; wherein, the target boxes and class labels of the target objects appended for annotation this time are obtained by inputting the first training image into the updated target detection model, and the similarity between the target objects appended for annotation this time and the annotated target objects is greater than the preset threshold;
[0183] The updated object detection model is obtained by training the object detection model based on the first training image that includes the object bounding boxes and class labels annotated by the user, as well as the object bounding boxes and class labels of the previously appended annotations.
[0184] That is to say, the user can issue the label append instruction multiple times until all the target objects on the first training image are annotated, or until the object detection model cannot detect new target objects. And before each response to the label append instruction, the electronic device can retrain the object detection model using the object bounding boxes and class labels annotated by the user, as well as the first training image with the object bounding boxes and class labels of the previously appended annotations, so as to obtain the updated object detection model.
[0185] For example, if there are 10 target objects in the first training image, the user can first manually annotate one target object in the first training image, and then issue the label append instruction. In response to this label append instruction, the electronic device appends the annotation of 3 target objects. Then the user can issue the label append instruction again. In response to this label append instruction, the electronic device appends the annotation of 2 target objects... and so on in a loop. Eventually, after the electronic device appends the annotation of all the target objects, the user can stop issuing the label append instruction.
[0186] Exemplarily, referring to FIG. 12(b), FIG. 12(b) is a schematic diagram showing that all the target objects are annotated in the image annotation method provided in the embodiment of the present application.
[0187] S1104, complete the intelligent append process.
[0188] In this embodiment, the electronic device can gradually release the result of the appended annotation, leaving enough amount for the user to gradually correct the object detection model, which can greatly improve the accuracy of the appended annotation process and reduce the detection error caused by external factors such as the environment.
[0189] For ease of understanding, referring to Figure 13 , Figure 13 is Figure 4 a schematic diagram of the intelligent append interaction principle of the illustrated embodiment. As shown in Figure 13 shown, this process includes the following steps:
[0190] S1301, the intelligent append function is default disabled; that is, in the initial interactive interface displayed by the electronic device, the intelligent append control is not unlocked.
[0191] S1302, at least one target has been marked in the current image. At this time, the condition for unlocking the intelligent append function is met, and then S1303 is executed;
[0192] S1303, the intelligent append function is released from the disabled state;
[0193] S1304, Click on "Intelligent Append"; that is, the user clicks on "Intelligent Append".
[0194] S1305, Invoke the algorithm, and based on the existing annotation content, mark the most similar targets (i.e., target objects) located on the image, with the number being 1 - 3.
[0195] This step is the process of training the object detection model using the previously annotated first training image and then inputting the first training image into the object detection model as described above.
[0196] S1306, Determine whether the appended content is correct; that is, the user determines whether the target objects intelligently appended are correct; if so, execute S1307; if not, execute S1308;
[0197] S1307, Determine whether all the targets in the current image have been appended; that is, the user determines whether all the targets in the current image have been appended; if so, execute S1309; if not, return to execute 1304;
[0198] S1308, Modify and adjust the marked content; that is, the user modifies and adjusts the target boxes and class labels of the marked target objects.
[0199] S1309, Switch to the next image.
[0200] In an embodiment of the present application, refer to Figure 14 , Figure 14 is Figure 4 the schematic diagram of the automatic annotation process in the embodiment shown, as Figure 14 shown, this process may include the following steps:
[0201] S1401, Display the first second training image to be annotated;
[0202] S1402, In response to the automatic labeling instruction, on the second training image, display the target boxes and class labels of the target objects annotated after the second training image is input into the object detection model; and, based on user selection, on the second training image, display the target boxes and class labels of all the target objects after the user's modification and adjustment.
[0203] After each execution of the automatic labeling instruction on a second training image, the image annotation method provided by the embodiment of the present application may further include:
[0204] S1403, Based on user selection, add the second training image modified and adjusted by the user to the target training set to update the target training set;
[0205] Specifically, after the user determines that the annotation of a second training image is completed, the user can decide whether to add the second training image to the target training set. If the user determines that the second training image needs to be added to the target training set, the electronic device can first determine whether the second training image has been manually corrected and adjusted by the user. If the second training image has not been corrected and adjusted, the electronic device can refuse to add the second training image to the target training set and display a text prompt to the user. For example, refer to Figure 15 , Figure 15 is the eighth schematic diagram of the interactive interface shown in Fig. 5(a) (the state of prompting to refuse to add an image to the target training set). As Figure 15 shown, text can be displayed below the second training image: "The image after automatic labeling will not participate in training. Add it to the training samples after manual adjustment."
[0206] If the second training image has been corrected and adjusted, the second training image can be added to the target training set.
[0207] S1404, display the next second training image to be annotated;
[0208] After each display of the next second training image to be annotated and before responding to the automatic labeling instruction, the annotation method of this image can further include:
[0209] S1405, based on the user's selection, train the target detection model with the updated target training set to update the target detection model;
[0210] That is to say, after each display of the next second training image, the user can decide whether to update the target detection model. If an update is needed, the electronic device can use all the images in the current target training set to update the target detection model; if no update is needed, the target detection model for automatically annotating the previous second training image can continue to be used.
[0211] For the second and subsequent second training images, in response to the automatic labeling instruction, the following is displayed on the second training image: the target box and class label of the target object in the second training image, including:
[0212] S1406, based on the user's selection, on the second training image, display the target box and class label of the target object annotated after the second training image is input into the unupdated or updated target detection model; and, based on the user's selection, on the second training image, display the target boxes and class labels of all target objects after the user's correction and adjustment.
[0213] Refer to Figure 16 , Figure 16 is the ninth schematic diagram of the interactive interface shown in Fig. 5(a). AsFigure 16 As shown, after each display of the next second training image, the electronic device can display a Toast prompt box and display two controls in the prompt box: "Automatic Labeling" and "Retrain and Automatic Labeling". Of course, this process is not limited to this display method, and this is only an example here. If the user clicks on "Automatic Labeling", it means that the target detection model does not need to be updated. At this time, the second training image is input into the unupdated target detection model. If the user clicks on "Retrain and Automatic Labeling", it means that the target detection model needs to be updated. At this time, the second training image is input into the updated target detection model.
[0214] For ease of understanding, refer to Figure 17 , Figure 17 is Figure 4 a schematic diagram of the automatic labeling interaction principle of the embodiment shown. As Figure 17 shown, this process may include the following steps:
[0215] S1701, the automatic labeling function is default disabled; that is, in the initial interactive interface displayed by the electronic device, the automatic labeling function is not unlocked;
[0216] S1702, content has been labeled on an image;
[0217] In the case of switching to an image without labels, the automatic labeling function unlocking condition is met, and S1703 is executed;
[0218] S1703, the automatic labeling function is released from the disabled state and relevant instructions are displayed in the interactive interface; for relevant instructions, refer to Figure 10 .
[0219] S1704, click on "Train and Automatic Labeling"; that is, the user clicks on "Train and Automatic Labeling";
[0220] S1705, mark all the inferred targets on the image; that is, the electronic device will train the model based on the existing annotation content, and then mark all the inferred target objects on the image.
[0221] When this step is executed for the first time, it is the process of training the target detection model using the previously labeled first training image and then inputting the second training image into the target detection model. When this step is executed for the second time and later, this step is the process of displaying the target boxes and class labels of the target objects marked after the second training image is input into the unupdated or updated target detection model based on the user's selection.
[0222] S1706, Determine whether to adjust the automatic annotation; that is, the user can decide whether to adjust the automatic annotation. If so, execute S1707; if not, execute S1708.
[0223] S1707, Adjust the marked content or supplement the unautomatically marked targets; that is, the user adjusts the target boxes and category labels of the marked target objects, or supplements the unautomatically marked target objects.
[0224] S1708, Confirm whether to include the current image in the target training set; that is, the user confirms whether to include the current image in the target training set. If so, execute S1707. At this time, the interaction interface will display a prompt message when the image has not been manually adjusted to prompt the user that the images included in the target training set need to be manually adjusted. If not, execute S1709;
[0225] S1709, Switch to the next unlabeled image;
[0226] S1710, Confirm whether to continue automatic labeling using the current model; that is, after switching to the next image, the interaction interface will display a prompt message for the user to confirm whether to continue automatic labeling using the current model, as Figure 16 shown. If so, execute S1711; if not, execute S1712;
[0227] S1711, Continue automatic labeling using the current model; at this time, the user can execute S1713: click "Automatic Labeling", and then the interaction logic will return to S1705 again;
[0228] S1712, Retrain a new model according to the current target training set and then perform automatic labeling; that is, update the target detection model and use the updated target detection model for automatic labeling; at this time, the user can execute S1714: click "Retrain and Automatic Labeling", and then the interaction logic will return to S1705 again. By looping like this, all images can be labeled.
[0229] In this embodiment, the user can screen the training images to determine whether to add the training images to the target training set. Furthermore, the user can also update the target detection model using the images in the target training set to continuously improve the performance of the target detection model and make the function of automatic annotation more and more accurate.
[0230] Next, a detailed description of the second method for annotating training images will be given.
[0231] This application embodiment also provides an image annotation method, which is applied to an electronic device. Refer to Figure 18 , Figure 18Schematic flowchart of the first embodiment of the second annotation method for images provided by the embodiments of the present application, as Figure 18 shown, the method includes:
[0232] S1801, display a training image to be annotated;
[0233] S1802, receive the annotation of the user for the training image, and display the target boxes and class labels of different target objects annotated by the user on the training image;
[0234] Among them, each class label annotates at least one target object;
[0235] S1803, in response to the label append instruction, display on the training image: the target boxes and class labels of the target objects in the training image;
[0236] Among them, the target boxes and class labels of the target objects in the training image are obtained by inputting the training image into a target detection model trained with the training image annotated by the user.
[0237] The above steps are similar to S403 - S406, and the similar parts will not be elaborated here. The difference between this embodiment and Figure 4 the embodiment shown is that in this embodiment, the annotation of the image can be completed through the intelligent append function without using the automatic annotation function.
[0238] In this embodiment, only some of the target objects in the training image need to be manually annotated, and the other target objects can be automatically annotated by the electronic device. Therefore, through this solution, the image can be annotated more efficiently, reducing the time cost of annotation.
[0239] See Figure 19 , Figure 19 Schematic flowchart of the second embodiment of the second annotation method for images provided by the embodiments of the present application, as Figure 19 shown, the method includes:
[0240] S1901, display a training image to be annotated;
[0241] S1902, receive the annotation of the user for the training image, and display the target boxes and class labels of different target objects annotated by the user on the training image; among them, each class label annotates at least one target object;
[0242] S1903, in response to the label append instruction, display on the training image: the target boxes and class labels of the target objects in the training image;
[0243] Among them, the target box and class label of the target object in the training image are obtained by inputting the training image into an object detection model trained with the training image annotated by the user;
[0244] The above steps are similar to S403 - S406, and the similar parts will not be elaborated here.
[0245] After step S1903, the above steps S1102 - S1103 can also be executed until the target boxes and class labels of all target objects are displayed on the current training image, and then step S1904 is executed.
[0246] S1904, determine whether there is still a training image to be annotated; if not, execute S1906; if so, execute S1905;
[0247] S1905, display the next training image to be annotated; then it can return to execute step S1902.
[0248] S1906, complete the process of intelligent addition.
[0249] That is to say, in the case where the number of training images to be annotated is multiple; after executing the label addition instruction for the first training image, the next training image to be annotated can be sequentially displayed. After displaying the next training image to be annotated, the training image can also be manually annotated by the user, and then the label addition instruction can be executed for this training image. And so on, until all training images are labeled.
[0250] In this embodiment, only some of the target objects in the training image need to be manually annotated by the user, and the other target objects in this training image can be automatically annotated by the electronic device. Therefore, through this solution, the image can be annotated more efficiently, reducing the time cost of annotation.
[0251] In one embodiment, in the second annotation method of the image provided in the embodiments of the present application, the user can also annotate the training image in an automatic annotation manner. For details, see Figure 20 , Figure 20 is the flowchart of the third embodiment of the second annotation method of the image provided in the embodiments of the present application. As Figure 20 shown, this method includes:
[0252] S2001, display a training image to be annotated;
[0253] S2002, receive the annotation of the user for this training image, and display the target boxes and class labels of different target objects annotated by the user on this training image; among them, at least one target object is annotated for each class label;
[0254] In S2003, in response to a label addition instruction, display on the training image: the target bounding box and class label of the target object in the training image;
[0255] Among them, the target bounding box and class label of the target object in the training image are obtained by inputting the training image into a target detection model trained with the training image annotated by the user;
[0256] The above steps are similar to S403 - S406, and the similar parts will not be elaborated here.
[0257] After step S2003, the above steps S1102 - S1103 can also be executed until the target bounding boxes and class labels of all target objects are displayed on the current training image, and then execute step S2004.
[0258] In S2004, determine whether there is still a training image to be annotated; if not, execute S2007; if so, execute S2005;
[0259] In S2005, display the next training image to be annotated;
[0260] In S2006, in response to an automatic labeling instruction, display on the next training image: the target bounding box and class label of the target object in the training image; among them, the target bounding box and class label of the target object in the training image are obtained by inputting the training image into a target detection model trained with the already annotated training images; then it can return to execute step S2004 until all training images are labeled.
[0261] In S2007, complete the process of automatic annotation.
[0262] Based on the above image annotation method, an embodiment of the present application also provides a model training method. Refer to Figure 21 , Figure 21 which is a schematic flowchart of the second embodiment of the model training method provided by the embodiment of the present application. As Figure 21 shown, the method includes the following steps:
[0263] In S2101, obtain the training image determined by the user;
[0264] In S2102, apply any of the above - mentioned image annotation methods to annotate the obtained training image;
[0265] In S2103, in response to a model training instruction, use the already annotated training images to train the target detection model to obtain a trained target detection model.
[0266] The model training instruction can also be issued by the user by clicking on the corresponding control in the interaction interface. For example, the control can be Figure 6 the "Train" button in
[0267] The embodiments of the present application do not limit the specific training method. Exemplarily, the training can be performed as follows:
[0268] First, input the unlabeled training images into the target detection model to be trained, and obtain the target boxes and class labels of the target objects detected by the model for the images as the model prediction results; calculate the model loss based on the model prediction results and the annotation results of the training images, and use the calculated model loss to adjust the parameters of the target detection model until the model converges, and obtain the trained target detection model.
[0269] In this embodiment, only some of the target objects in the training images need to be manually annotated, and the remaining target objects in the training images and the target objects in other training images can be automatically annotated by the electronic device. Therefore, through this solution, the images can be annotated more efficiently, the time cost of annotation can be reduced, and the training efficiency of the target detection model can be improved accordingly.
[0270] The embodiments of the present application also provide an image annotation device, which is applied to an electronic device. Refer to Figure 22 , Figure 22 which is the structural schematic diagram of the first embodiment of the image annotation device provided by the embodiments of the present application; as Figure 22 shown, the device includes:
[0271] A display module 2201, configured to display a first training image to be annotated;
[0272] An annotation receiving module 2202, configured to receive the annotation of the user for the first training image, and display the target boxes and class labels of different target objects annotated by the user on the first training image; wherein, at least one target object is annotated with each class label;
[0273] An automatic labeling response module 2203, configured to, in response to an automatic labeling instruction, display on a second training image: the target boxes and class labels of the target objects in the second training image; wherein, the target boxes and class labels of the target objects in the second training image are obtained by inputting the second training image into a target detection model trained using the labeled first training image; the labeled first training image includes: the target boxes and class labels annotated by the user.
[0274] In one embodiment, the device further includes: a label append response module, configured to, after the annotation receiving module 2202 receives the user's annotation for the first training image and displays the target boxes and class labels of different target objects annotated by the user on the first training image, in response to a label append instruction, display on the first training image: the target boxes and class labels of the target objects in the first training image; wherein, the target boxes and class labels of the target objects in the first training image are obtained by inputting the first training image into a target detection model trained using the first training image annotated by the user; the annotated first training image further includes: the target boxes and class labels annotated after being appended.
[0275] In one embodiment, the device further includes:
[0276] a label creation response module, configured to, before the annotation receiving module 2202 receives the user's annotation for the first training image, in response to a label creation instruction, receive and save different class labels corresponding to different target objects created by the user for the first training image.
[0277] In one embodiment, the device further includes: an image switching module, configured to, when the number of second training images to be annotated is greater than 1, after the automatic labeling response module 2203 finishes executing the automatic labeling instruction on the first second training image, sequentially display the next second training image to be annotated, and execute the automatic labeling instruction for this second training image until all second training images are labeled.
[0278] In one embodiment, the device further includes: a training atlas acquisition module, configured to obtain a training atlas determined by the user before the display module 2201 displays the first training image to be annotated; wherein, the training atlas includes multiple training images;
[0279] a training image receiving module, configured to receive a training image selected by the user from the training atlas as the first training image to be annotated; wherein, except for this first training image to be annotated in the training atlas, other images are used as the second training images to be annotated.
[0280] In one embodiment, the automatic labeling response module 2203 is specifically configured to:
[0281] For the first second training image, in response to the automatic labeling instruction, display on the second training image the target boxes and class labels of the target objects annotated after the second training image is input into the target detection model; and, based on the user's selection, display on the second training image the target boxes and class labels of all target objects after being corrected and adjusted by the user.
[0282] In one embodiment, the device further includes: a target training set updating module, configured to, after each execution of the automatic tagging instruction on a second training image, based on user selection, add the second training image that has been corrected and adjusted by the user to the target training set to update the target training set;
[0283] a model updating module, configured to, after each time the image switching module displays the next second training image to be annotated and before the automatic tagging response module 2203 responds to the automatic tagging instruction, based on user selection, train the target detection model with the updated target training set to update the target detection model;
[0284] The automatic tagging response module 2203 is specifically configured to:
[0285] Based on user selection, on the second training image, display the target boxes and class labels of the target objects annotated after the second training image is input into the unupdated or updated target detection model; and, based on user selection, on the second training image, display the target boxes and class labels of all target objects after the user's correction and adjustment.
[0286] In one embodiment, the label append response module is specifically configured to:
[0287] In response to the label append instruction, display the target boxes and class labels of the target objects newly annotated this time on the first training image; wherein, the target boxes and class labels of the target objects newly annotated this time are obtained by inputting the first training image into the target detection model trained with the first training image annotated by the user, and the similarity between the target objects newly annotated this time and the target objects annotated by the user is greater than a preset threshold.
[0288] In one embodiment, the label append response module is further configured to, after the label append instruction is executed and when the target boxes and class labels of all target objects are not displayed on the first training image, in response to the next label append instruction, display the target boxes and class labels of the target objects newly annotated this time on the first training image; wherein, the target boxes and class labels of the target objects newly annotated this time are obtained by inputting the first training image into the updated target detection model, and the similarity between the target objects newly annotated this time and the already annotated target objects is greater than a preset threshold; the updated target detection model is obtained by training the target detection model based on the first training image including the target boxes and class labels annotated by the user and the target boxes and class labels newly annotated before.
[0289] In one embodiment, the device further includes:
[0290] A correction and adjustment module, configured to, after the label appending response module displays the target box and category label of the target object for the current appending annotation, correct and adjust the target box and category label of the target object for the current appending annotation based on user selection.
[0291] In one embodiment, the apparatus further includes:
[0292] A first control display module, configured to display a label creation control before the label creation response module responds to a label creation instruction; wherein, the label creation instruction is generated by the user clicking the label creation control;
[0293] A second control display module, configured to display label controls corresponding to different target objects created by the user before the annotation receiving module 2202 receives the annotation of the user for the first training image; wherein, the target box and category label of each target object annotated by the user on the first training image are obtained by annotation when the user selects the label control corresponding to the target object.
[0294] A third control display module, configured to display an automatic labeling control before the automatic labeling response module 2203 responds to an automatic labeling instruction; wherein, the automatic labeling instruction is generated by the user clicking the automatic labeling control;
[0295] A fourth control display module, configured to display an intelligent appending control before the label appending response module responds to a label appending instruction; wherein, the label appending instruction is generated by the user clicking the intelligent appending control.
[0296] In one embodiment, the automatic labeling response module 2203 is specifically configured to: respond to an automatic labeling instruction when the automatic labeling control is unlocked; wherein, the unlocking condition of the automatic labeling control is: the first training image has been annotated, and a second training image to be annotated is displayed;
[0297] The label appending response module is specifically configured to: respond to a label appending instruction when the intelligent appending control is unlocked; wherein, the unlocking condition of the intelligent appending control is: there are target boxes and category labels annotated by the user on the first training image.
[0298] In one embodiment, the apparatus further includes:
[0299] A modification response module, configured to, when receiving a modification instruction for any target box, modify the position, size, and / or corresponding category label of the target box based on the modification instruction;
[0300] A deletion response module, configured to remove the target box and the corresponding category label from the training image when a deletion instruction for any target box is received;
[0301] An annotation addition module, configured to display the target boxes and category labels newly added by the user on the training image when an annotation newly added by the user for the training image is received.
[0302] An embodiment of the present application further provides an image annotation device, which is applied to an electronic device. Refer to Figure 23 , Figure 23 is a schematic structural diagram of a second embodiment of the image annotation device provided by the embodiment of the present application; the device includes:
[0303] A training image display module 2301, configured to display a training image to be annotated;
[0304] A user annotation receiving module 2302, configured to receive an annotation of the user for the training image, and display the target boxes and category labels of different target objects annotated by the user on the training image; wherein, at least one target object is annotated with each category label;
[0305] A label append instruction response module 2303, configured to, in response to a label append instruction, display on the training image: the target boxes and category labels of the target objects in the training image; wherein, the target boxes and category labels of the target objects in the training image are obtained by inputting the training image into a target detection model trained with the training image annotated by the user.
[0306] In one embodiment, the number of training images to be annotated is multiple; the device further includes:
[0307] An image append module, configured to, after executing the label append instruction for the first training image, sequentially display the next training image to be annotated, and execute the label append instruction for this training image until all training images are labeled.
[0308] In one embodiment, the number of training images to be annotated is multiple; the device further includes:
[0309] An automatic labeling module, configured to, after executing the label append instruction for the first training image, sequentially display the next training image to be annotated; in response to an automatic labeling instruction, display on the next training image: the target boxes and category labels of the target objects in this training image; wherein, the target boxes and category labels of the target objects in this training image are obtained by inputting this training image into a target detection model trained with the already annotated training image; return to execute the step of sequentially displaying the next training image to be annotated until all training images are labeled.
[0310] An embodiment of the present application further provides a model training device, which is applied to an electronic device. Refer to Figure 24 , Figure 24 which is a schematic structural diagram of the model training device provided by the embodiment of the present application. As Figure 24 shown, the device includes:
[0311] A training image acquisition module 2401, configured to acquire training images determined by a user;
[0312] A training image annotation module 2402, configured to annotate the acquired training images by using any of the above-mentioned image annotation methods;
[0313] A model training instruction response module 2403, configured to respond to a model training instruction, and train a target detection model by using the annotated training images to obtain a trained target detection model.
[0314] An embodiment of the present application further provides an electronic device. Refer to Figure 25 , Figure 25 which is a schematic structural diagram of the electronic device provided by the embodiment of the present application. As Figure 25 shown, it includes:
[0315] A memory 2501, configured to store a computer program;
[0316] A processor 2502, configured to implement the steps of any of the above-mentioned image annotation methods or model training methods when executing the program stored in the memory 2501.
[0317] Moreover, the above-mentioned electronic device may further include a communication bus and / or a communication interface, and the processor 2502, the communication interface, and the memory 2501 complete communication with each other through the communication bus.
[0318] The communication bus mentioned in the above-mentioned electronic device may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus may be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity, only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.
[0319] The communication interface is used for communication between the above-mentioned electronic device and other devices.
[0320] The memory may include a Random Access Memory (RAM), or may also include a Non-Volatile Memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located away from the aforementioned processor.
[0321] The aforementioned processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0322] In another embodiment provided by the present application, a computer-readable storage medium is further provided. A computer program is stored in the computer-readable storage medium, and when the computer program is executed by a processor, the steps of any of the above-mentioned image annotation methods or model training methods are implemented.
[0323] In another embodiment provided by the present application, a computer program product containing instructions is further provided. When it runs on a computer, the computer is caused to execute any of the image annotation methods or model training methods in the above embodiments.
[0324] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from a website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that a computer can access, or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a solid-state disk (SSD), etc.
[0325] It should be noted that in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including", or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article, or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article, or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article, or device including the element.
[0326] Each embodiment in this specification is described in a related manner. The same or similar parts among the embodiments can be referred to each other, and the differences between each embodiment and other embodiments are emphasized. In particular, for the embodiments of the device, electronic device, readable storage medium, and computer program product, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments.
[0327] The above are only the preferred embodiments of the present application and are not intended to limit the protection scope of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application are all included in the protection scope of the present application.
Claims
1. A method for labeling an image, characterized in that: Applied to electronic equipment, the method comprises: displaying a first training image to be labeled; receiving annotations of the first training image by a user, and displaying target frames and category labels of different target objects annotated by the user on the first training image; wherein each category label annotates at least one target object; In response to the automatic labeling instruction, the target frame and category label of the target object in the second training image are displayed on the second training image; wherein the target frame and category label of the target object in the second training image are obtained by inputting the second training image into a target detection model trained using the labeled first training image; the labeled first training image includes: the target frame and category label labeled by the user.
2. The method according to claim 1, characterized in that After receiving the annotations of the user on the first training image and displaying the target boxes and category labels of different target objects annotated by the user on the first training image, the method further includes: In response to the label appending instruction, displaying on the first training image: a target frame and a category label of the target object in the first training image; The target frame and category label of the target object in the first training image are obtained by inputting the first training image into a target detection model trained using the first training image annotated by the user; The labeled first training image further includes: a target frame and a category label that are additionally labeled.
3. The method according to claim 2, characterized in that Before receiving the annotation of the first training image by the user, the method further includes: In response to the label creation instruction, different category labels corresponding to different target objects created by the user for the first training image are received and saved.
4. The method according to claim 1 or 2, characterized in that: When the number of the second training images to be labeled is greater than 1, after the automatic labeling instruction is executed on the first second training image, the method further includes: The next second training image to be labeled is displayed in sequence, and the automatic labeling instruction is executed for the second training image until the labeling of all the second training images is completed.
5. The method according to claim 4, characterized in that Before the step of displaying the first training image to be labeled, the method further includes: Obtaining a training atlas determined by a user; wherein the training atlas includes a plurality of training images; A training image selected by a user from the training atlas is received as a first training image to be labeled; wherein, except for the first training image to be labeled, other images in the training atlas are used as second training images to be labeled.
6. The method according to claim 4, characterized in that For the first second training image, in response to the automatic labeling instruction, displaying on the second training image: a target frame and a category label of a target object in the second training image, including: In response to the automatic labeling instruction, displaying, on the second training image, a target frame and a category label of the target object annotated after the second training image is input into the target detection model; and Based on the user selection, the target boxes and category labels of all target objects corrected and adjusted by the user are displayed on the second training image.
7. The method according to claim 6, characterized in that After executing the automatic marking instruction on a second training image each time, the method further includes: adding the second training image modified and adjusted by the user to the target training set based on the user's selection, so as to update the target training set; After displaying the next second training image to be labeled each time and before responding to the automatic labeling instruction, the method further includes: training the target detection model with the updated target training set based on the user selection to update the target detection model; For the second and subsequent second training images, in response to the automatic labeling instruction, displaying on the second training image: a target frame and a category label of the target object in the second training image, including: Based on the user's selection, displaying on the second training image the target box and the category label of the target object annotated after the second training image is input into the unupdated or updated target detection model; and Based on the user selection, the target boxes and category labels of all target objects corrected and adjusted by the user are displayed on the second training image.
8. The method according to claim 2, characterized in that: In response to the label appending instruction, displaying on the first training image: a target frame and a category label of a target object in the first training image, including: In response to the label addition instruction, the target frame and category label of the target object additionally annotated this time are displayed on the first training image; wherein, the target frame and category label of the target object additionally annotated this time are obtained by inputting the first training image into a target detection model trained using the first training image annotated by the user, and the similarity between the target object additionally annotated this time and the target object annotated by the user is greater than a preset threshold.
9. The method according to claim 8, characterized in that After the label appending instruction is executed, if the target frames and category labels of all target objects are not displayed on the first training image, the method further includes: In response to the next label addition instruction, displaying the target frame and category label of the target object additionally annotated this time on the first training image; wherein the target frame and category label of the target object additionally annotated this time are obtained by inputting the first training image into the updated target detection model, and the similarity between the target object additionally annotated this time and the annotated target object is greater than a preset threshold; The updated target detection model is obtained by training the target detection model based on the first training image including the target frame and category label annotated by the user and the target frame and category label annotated previously.
10. The method according to claim 9, characterized in that After displaying the target frame and category label of the target object additionally annotated this time, the method further includes: Based on the user's selection, the target box and category label of the target object to be additionally annotated are corrected and adjusted.
11. The method according to claim 3, characterized in that Before responding to the label creation instruction, the method further includes: displaying a label creation control; wherein the label creation instruction is generated by the user by clicking the label creation control; Before receiving the user's annotations for the first training image, the method further includes: displaying label controls corresponding to different target objects created by the user; wherein the target box and category label of each target object annotated by the user on the first training image are obtained by annotating when the user selects the label control corresponding to the target object; Before responding to the automatic marking instruction, the method further includes: displaying an automatic marking control; wherein the automatic marking instruction is generated by the user by clicking the automatic marking control; Before responding to the tag-appending instruction, the method further includes: displaying a smart-appending control; wherein the tag-appending instruction is generated by the user by clicking the smart-appending control.
12. The method according to claim 11, characterized in that The responding to the automatic marking instruction includes: responding to the automatic marking instruction when the automatic marking control is unlocked; wherein the unlocking condition of the automatic marking control is: the first training image has been marked, and the second training image to be marked is displayed; The responding to the label appending instruction includes: responding to the label appending instruction when the smart appending control is unlocked; wherein the unlocking condition of the smart appending control is: there are a target box and a category label marked by the user on the first training image.
13. The method according to claim 1, characterized in that The method further comprises: When a modification instruction for any target frame is received, modifying the position, size, and / or corresponding category label of the target frame based on the modification instruction; When a deletion instruction is received for any target box, the target box and the corresponding category label are removed from the training image; When a new annotation added by the user for a training image is received, the target box and category label added by the user are displayed on the training image.
14. A method for labeling an image, characterized in that: Applied to electronic equipment, the method comprises: Display a training image to be labeled; receiving annotations of the training image by a user, and displaying target frames and category labels of different target objects annotated by the user on the training image; wherein each category label annotates at least one target object; In response to the label appending instruction, displaying on the training image: a target frame and a category label of a target object in the training image; The target frame and category label of the target object in the training image are obtained by inputting the training image into a target detection model trained using the training image annotated by the user.
15. The method according to claim 14, characterized in that The number of the training images to be labeled is multiple; after executing the label appending instruction for the first training image, the method further includes: The next training image to be labeled is displayed in sequence, and the label appending instruction is executed for the training image until all training images are labeled.
16. The method according to claim 14, characterized in that After executing the label appending instruction for the first training image, the method further includes: Display the next training image to be labeled in sequence; In response to the automatic labeling instruction, displaying on the next training image: a target frame and a category label of a target object in the training image; wherein the target frame and the category label of the target object in the training image are obtained by inputting the training image into a target detection model trained using the labeled training image; Return to the step of sequentially displaying the next training image to be labeled until all training images are labeled.
17. A model training method, characterized in that: Applied to electronic equipment, the method comprises: Obtaining a training image determined by a user; Applying the image annotation method according to any one of claims 1 to 16 to annotate the training image; In response to the model training instruction, the target detection model is trained using the labeled training images to obtain a trained target detection model.
18. An image annotation device, characterized in that: Applied to electronic equipment, the device comprises: A display module, used for displaying a first training image to be labeled; a label receiving module, configured to receive a user's labeling for the first training image, and display target frames and category labels of different target objects labeled by the user on the first training image; wherein each category label labels at least one target object; The automatic labeling response module is used to respond to the automatic labeling instruction and display on the second training image: the target frame and category label of the target object in the second training image; wherein the target frame and category label of the target object in the second training image are obtained by inputting the second training image into the target detection model trained by the labeled first training image; the labeled first training image includes: the target frame and category label labeled by the user.
19. An image annotation device, characterized in that: Applied to electronic equipment, the device comprises: A training image display module, used to display a training image to be labeled; A user annotation receiving module, configured to receive user annotations for the training image, and display target frames and category labels of different target objects annotated by the user on the training image; wherein each category label annotates at least one target object; The label appending instruction response module is used to display on the training image: the target frame and category label of the target object in the training image; wherein the target frame and category label of the target object in the training image are obtained by inputting the training image into a target detection model trained using the training image annotated by the user.
20. A model training device, characterized in that: Applied to electronic equipment, the device comprises: A training image acquisition module, used to obtain a training image determined by a user; A training image annotation module, used to apply the image annotation method according to any one of claims 1 to 16 to annotate the training image; The model training instruction response module is used to respond to the model training instruction and use the labeled training images to train the target detection model to obtain a trained target detection model.
21. An electronic device, characterized in that: include: Memory, used to store computer programs; A processor, for implementing the method described in any one of claims 1 to 13, any one of claims 14 to 16, or claim 17 when executing a program stored in a memory.
22. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method of any one of claims 1 to 13, any one of claims 14 to 16, or claim 17 is implemented.
23. A computer program product comprising instructions, characterized in that When the computer program product is run on a computer, the computer is enabled to execute the method of any one of claims 1 to 13, any one of claims 14 to 16, or claim 17.
Citation Information
Patent Citations
Picture labeling method and device, storage medium and electronic device
CN110610169A
Data annotation method and device, electronic equipment and storage medium
CN111680753A
Data annotation model training method and device thereof, electronic equipment and storage medium
CN114078578A
Image labeling method and device, equipment and storage medium
CN117253104A
Image annotation model training method, image annotation method and related equipment
CN118038126A