Image annotation method, model training method, device and electronic equipment

By displaying images on electronic devices and receiving user annotations, combined with the automatic labeling and intelligent appending functions of the target detection model, the problem of low image annotation efficiency is solved and efficient automatic image annotation is achieved.

CN120126141BActive Publication Date: 2025-09-05HANGZHOU HIKROBOT TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510590431.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-08
Publication Date
2025-09-05
Estimated Expiration
2045-05-08

AI Technical Summary

Technical Problem

The existing image annotation technology has low efficiency, manual annotation is time-consuming and labor-intensive, and it is difficult to efficiently complete the annotation of multiple target objects in multiple images.

Method used

By displaying the image to be labeled on an electronic device, receiving the target box and category label labeled by the user, and using the target detection model trained with the labeled image to automatically label or additionally label unlabeled targets, combined with the intelligent append function, the number of manual labeling times is reduced.

Benefits of technology

It improves the efficiency of image annotation, reduces time cost, and realizes efficient automatic image annotation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120126141B_ABST
    Figure CN120126141B_ABST
Patent Text Reader

Abstract

The present application provides an image annotation method, a model training method, an apparatus, and an electronic device, relating to the field of machine vision technology. The image annotation method, applied to an electronic device, includes: displaying a first training image to be annotated; receiving user annotations for the first training image, and displaying target boxes and category labels for different target objects annotated by the user on the first training image; and displaying, in response to an automatic labeling instruction or a label appending instruction, on a second training image: target boxes and category labels for the target objects in the second training image. This solution can improve image annotation efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of machine vision technology, and in particular to image annotation methods, model training methods, devices, and electronic equipment. Background Art

[0002] For target detection models, image annotation refers to boxing out the location of the target object of interest in the image and adding a category label representing its category, so that when training the target detection model, the model can accurately extract the features of the target object.

[0003] However, currently, image labeling is usually done manually. During a training process of a target detection model, multiple images usually need to be labeled, and each image may contain multiple target objects. Therefore, a large amount of labeling is required, which makes the labeling inefficient and consumes a lot of time. Summary of the Invention

[0004] The purpose of the embodiments of the present application is to provide an image annotation method, a model training method, an apparatus, and an electronic device to improve the efficiency of image annotation. The specific technical solutions are as follows:

[0005] In a first aspect, an embodiment of the present application provides an image annotation method, applied to an electronic device, the method comprising:

[0006] displaying a first training image to be labeled;

[0007] receiving annotations for the first training image from a user, and displaying target frames and category labels of different target objects annotated by the user on the first training image; wherein each category label annotates at least one target object;

[0008] In response to the automatic labeling instruction, the target box and category label of the target object in the second training image are displayed on the second training image; wherein the target box and category label of the target object in the second training image are obtained by inputting the second training image into the target detection model trained using the labeled first training image; the labeled first training image includes the target box and category label labeled by the user.

[0009] In one embodiment, after receiving user annotations for the first training image and displaying target boxes and category labels of different target objects annotated by the user on the first training image, the method further includes:

[0010] In response to a label appending instruction, displaying on the first training image: a target frame and a category label of a target object in the first training image;

[0011] The target box and category label of the target object in the first training image are obtained by inputting the first training image into the target detection model trained using the first training image annotated by the user;

[0012] The labeled first training image further includes: an additionally labeled target frame and category label.

[0013] In one embodiment, before receiving the user's annotations for the first training image, the method further includes:

[0014] In response to the label creation instruction, different category labels corresponding to different target objects created by the user for the first training image are received and saved.

[0015] In one embodiment, when the number of second training images to be labeled is greater than one, after executing the automatic labeling instruction on the first second training image, the method further includes:

[0016] The next second training image to be labeled is displayed in sequence, and the automatic labeling instruction is executed for the second training image until the labeling of all second training images is completed.

[0017] In one embodiment, before the step of displaying the first training image to be labeled, the method further includes:

[0018] Obtaining a training atlas determined by a user; wherein the training atlas includes a plurality of training images;

[0019] A training image selected by a user from the training atlas is received as a first training image to be labeled; wherein, except the first training image to be labeled, other images in the training atlas are used as second training images to be labeled.

[0020] In one embodiment, for the first second training image, in response to the automatic labeling instruction, displaying on the second training image: a target frame and a category label of a target object in the second training image includes:

[0021] In response to the automatic labeling instruction, displaying on the second training image the target box and category label of the target object annotated after the second training image is input into the target detection model; and

[0022] Based on the user selection, the target boxes and category labels of all target objects after user correction and adjustment are displayed on the second training image.

[0023] In one embodiment, after executing the automatic marking instruction on a second training image each time, the method further includes: adding the second training image that has been modified and adjusted by the user to the target training set based on a user selection to update the target training set;

[0024] After each display of the next second training image to be labeled and before responding to the automatic labeling instruction, the method further includes: training the target detection model with the updated target training set based on user selection to update the target detection model;

[0025] For the second and subsequent second training images, in response to the automatic labeling instruction, displaying on the second training image: a target frame and a category label of the target object in the second training image, including:

[0026] Based on the user's selection, displaying on the second training image the target box and category label of the target object annotated after the second training image is input into the unupdated or updated target detection model; and

[0027] Based on the user selection, the target boxes and category labels of all target objects after user correction and adjustment are displayed on the second training image.

[0028] In one embodiment, in response to the label appending instruction, displaying on the first training image: a target frame and a category label of the target object in the first training image includes:

[0029] In response to the label appending instruction, the target frame and category label of the target object additionally annotated this time are displayed on the first training image; wherein, the target frame and category label of the target object additionally annotated this time are obtained by inputting the first training image into the target detection model trained using the first training image annotated by the user, and the similarity between the target object additionally annotated this time and the target object annotated by the user is greater than a preset threshold.

[0030] In one embodiment, after the label appending instruction is executed, if target frames and category labels of all target objects are not displayed on the first training image, the method further includes:

[0031] In response to the next label addition instruction, displaying a target bounding box and a category label of the newly annotated target object on the first training image; wherein the target bounding box and the category label of the newly annotated target object are obtained by inputting the first training image into the updated target detection model, and the similarity between the newly annotated target object and the previously annotated target object is greater than a preset threshold;

[0032] The updated target detection model is obtained by training the target detection model based on the first training image including the target frame and category label annotated by the user and the target frame and category label annotated previously.

[0033] In one embodiment, after displaying the target frame and category label of the target object that is additionally annotated, the method further includes:

[0034] Based on the user's selection, the target box and category label of the target object to be annotated are corrected and adjusted.

[0035] In one embodiment, before responding to the tag creation instruction, the method further includes: displaying a tag creation control; wherein the tag creation instruction is generated by the user by clicking the tag creation control;

[0036] Before receiving the user's annotations for the first training image, the method further includes: displaying label controls corresponding to different target objects created by the user; wherein the target box and category label of each target object annotated by the user on the first training image are obtained by the user selecting the label control corresponding to the target object;

[0037] Before responding to the automatic marking instruction, the method further includes: displaying an automatic marking control; wherein the automatic marking instruction is generated by the user by clicking the automatic marking control;

[0038] Before responding to the tag appending instruction, the method further includes: displaying a smart appending control; wherein the tag appending instruction is generated by the user by clicking the smart appending control.

[0039] In one embodiment, responding to the automatic marking instruction includes: responding to the automatic marking instruction when the automatic marking control is unlocked; wherein the unlocking condition of the automatic marking control is: the first training image has been marked and the second training image to be marked is displayed;

[0040] The responding to the label appending instruction includes: responding to the label appending instruction when the smart appending control is unlocked; wherein the unlocking condition of the smart appending control is: there are a target box and a category label marked by the user on the first training image.

[0041] In one embodiment, the method further includes:

[0042] When a modification instruction for any target frame is received, modifying the position, size, and / or corresponding category label of the target frame based on the modification instruction;

[0043] When a deletion instruction is received for any target box, the target box and the corresponding category label are removed from the training image;

[0044] When a new annotation added by the user for a training image is received, the target box and category label added by the user are displayed on the training image.

[0045] In a second aspect, an embodiment of the present application provides an image annotation method, applied to an electronic device, the method comprising:

[0046] Display a training image to be labeled;

[0047] receiving annotations for the training image from a user, and displaying target frames and category labels of different target objects annotated by the user on the training image; wherein each category label annotates at least one target object;

[0048] In response to a label appending instruction, displaying on the training image: a target frame and a category label of a target object in the training image;

[0049] The target frame and category label of the target object in the training image are obtained by inputting the training image into a target detection model trained using training images annotated by the user.

[0050] In one embodiment, there are multiple training images to be labeled; after executing the label appending instruction for the first training image, the method further includes:

[0051] The next training image to be labeled is displayed in sequence, and the label appending instruction is executed for the training image until all training images are labeled.

[0052] In one embodiment, after executing the label appending instruction for the first training image, the method further includes:

[0053] Display the next training image to be labeled in sequence;

[0054] In response to the automatic labeling instruction, displaying on the next training image: a target bounding box and a category label of a target object in the training image; wherein the target bounding box and the category label of the target object in the training image are obtained by inputting the training image into an object detection model trained using the labeled training image;

[0055] Return to the step of sequentially displaying the next training image to be labeled until all training images are labeled.

[0056] In a third aspect, an embodiment of the present application provides a model training method, applied to an electronic device, the method comprising:

[0057] Obtaining a user-determined training image;

[0058] Apply any of the above-mentioned image annotation methods to annotate the training image;

[0059] In response to the model training instruction, the target detection model is trained using the labeled training images to obtain a trained target detection model.

[0060] In a fourth aspect, an embodiment of the present application provides an image annotation device, which is applied to an electronic device, and the device includes:

[0061] A display module, configured to display a first training image to be labeled;

[0062] a label receiving module, configured to receive user labels for the first training image, and display target boxes and category labels of different target objects labeled by the user on the first training image; wherein each category label labels at least one target object;

[0063] The automatic labeling response module is configured to display, in response to an automatic labeling instruction, on a second training image: a target box and a category label of a target object in the second training image; wherein the target box and category label of the target object in the second training image are obtained by inputting the second training image into a target detection model trained using an annotated first training image; the annotated first training image includes the target box and category label annotated by the user.

[0064] In a fifth aspect, an embodiment of the present application provides an image annotation device, which is applied to an electronic device, and the device includes:

[0065] A training image display module, used to display a training image to be labeled;

[0066] A user annotation receiving module is configured to receive user annotations for the training image and display target frames and category labels of different target objects annotated by the user on the training image; wherein each category label annotates at least one target object;

[0067] The label appending instruction response module is used to display on the training image: the target box and category label of the target object in the training image; wherein, the target box and category label of the target object in the training image are obtained by inputting the training image into a target detection model trained using the training image annotated by the user.

[0068] In a sixth aspect, an embodiment of the present application provides a model training device, applied to an electronic device, the device comprising:

[0069] A training image acquisition module, used to obtain a training image determined by a user;

[0070] A training image annotation module, configured to apply any of the above-mentioned image annotation methods to annotate the training image;

[0071] The model training instruction response module is used to respond to the model training instruction and use the labeled training images to train the target detection model to obtain a trained target detection model.

[0072] In a seventh aspect, an embodiment of the present application provides an electronic device, including:

[0073] Memory for storing computer programs;

[0074] The processor is configured to implement any of the above-mentioned image annotation methods or model training methods when executing the program stored in the memory.

[0075] In an eighth aspect, an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the computer program implements any of the above-mentioned image annotation methods or model training methods.

[0076] In a ninth aspect, an embodiment of the present application further provides a computer program product comprising instructions, which, when executed on a computer, enables the computer to execute any of the above-described image annotation methods or model training methods.

[0077] Beneficial effects of the embodiments of the present application:

[0078] By applying the embodiment of the present invention, a user can first manually annotate the first training image displayed by the electronic device, so that the electronic device can use the annotated first training image to train the target detection model. In this way, the user can then use the automatic labeling instruction to instruct the electronic device to input the second training image into the trained target detection model, thereby automatically obtaining the target box and category label of the target object in the second training image. It can be seen that by applying the embodiment of the present invention, only the first training image needs to be manually annotated, and the second training image can be automatically annotated by the electronic device. Therefore, this solution can more efficiently annotate images and reduce the time cost of annotation.

[0079] Of course, it is not necessary to achieve all the advantages described above at the same time when implementing any product or method of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0080] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other embodiments can also be obtained based on these drawings.

[0081] Figure 1 Schematic diagram of the structure of the few-shot target detection training operator in the embodiment of the present application;

[0082] Figure 2 To utilize Figure 1 The schematic diagram of the principle flow of training the target detection model using the few-shot target detection training operator shown in FIG.

[0083] Figure 3 A schematic flow chart of a first embodiment of a first image annotation method provided in an embodiment of the present application;

[0084] Figure 4 A schematic flow chart of a second embodiment of the first image annotation method provided in an embodiment of the present application;

[0085] FIG5( a ) is a first schematic diagram of an interactive interface used in the image annotation method provided in an embodiment of the present application (in a state where no image is imported);

[0086] Figure 5(b) is a second schematic diagram of the interactive interface shown in Figure 5(a) (displaying the status of common parameters for target detection);

[0087] Figure 6 This is the third schematic diagram of the interactive interface shown in Figure 5(a) (the state of importing images);

[0088] Figure 7 This is the fourth schematic diagram of the interactive interface shown in Figure 5(a) (the state when creating a tag);

[0089] Figure 8 This is the fifth schematic diagram of the interactive interface shown in Figure 5(a) (the state after the label creation is completed);

[0090] Figure 9 This is a sixth schematic diagram of the interactive interface shown in FIG5(a) (the state after the user has annotated the first training image);

[0091] Figure 10 This is the seventh schematic diagram of the interactive interface shown in Figure 5(a) (showing the status of the automatic marking button);

[0092] Figure 11 for Figure 4 A schematic diagram of the process of intelligent appending in the embodiment shown;

[0093] FIG12( a ) is a schematic diagram of label addition in the image annotation method provided in an embodiment of the present application;

[0094] FIG12( b ) is a schematic diagram showing that all target objects are marked in the image marking method provided in an embodiment of the present application;

[0095] Figure 13 for Figure 4 Schematic diagram of the intelligent append interaction principle of the illustrated embodiment;

[0096] Figure 14 for Figure 4 A schematic diagram of the automatic marking process in the embodiment shown;

[0097] Figure 15 This is the eighth schematic diagram of the interactive interface shown in Figure 5(a) (the state in which the image is refused to be added to the target training set);

[0098] Figure 16 This is the ninth schematic diagram of the interactive interface shown in Figure 5(a) (the status prompting whether to retrain the model);

[0099] Figure 17 for Figure 4 Schematic diagram of the automatic marking interaction principle of the illustrated embodiment;

[0100] Figure 18 A schematic flow chart of a first embodiment of the second image annotation method provided in an embodiment of the present application;

[0101] Figure 19 A schematic flow chart of a second embodiment of the second image annotation method provided in an embodiment of the present application;

[0102] Figure 20 A schematic flow chart of a third embodiment of the second image annotation method provided in an embodiment of the present application;

[0103] Figure 21 A flowchart of an embodiment of the model training method provided in the embodiments of the present application;

[0104] Figure 22 A schematic structural diagram of a first embodiment of an image annotation device provided in an embodiment of the present application;

[0105] Figure 23 A schematic structural diagram of a second embodiment of the image annotation device provided in an embodiment of the present application;

[0106] Figure 24 A schematic diagram of the structure of the model training device provided in an embodiment of the present application;

[0107] Figure 25 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0108] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field based on this application are within the scope of protection of this application.

[0109] In this embodiment, the object detection model to be trained can be a deep learning model, such as a Few-Shot Object Detection (FSOD) model. In order to train this model, a dedicated training program can be developed. This program can be called a Few-Shot Object Detection Training Operator. Figure 1 , Figure 1 Schematic diagram of the structure of the few-sample target detection training operator in the embodiment of the present application; Figure 1 As shown, the program can encapsulate the following content: a basic deep learning model for target detection (i.e., the target detection model that needs to be trained, hereinafter referred to as the target detection model), common target detection parameters, and target detection annotation functionality. For example, common target detection parameters may include: angle enable, which allows the user to determine whether to use a rectangular box with an angle value to annotate the target object; optimal model resolution setting enable, which allows the user to determine whether to adjust the resolution of the displayed training image; and optimal model resolution, which allows the user to set the resolution of the displayed training image.

[0110] This program allows users to configure common parameters for object detection models and use the object detection annotation feature to mark the target objects of interest in training images. Finally, a basic object detection deep learning model is trained based on the annotated training images to obtain an object detection model suitable for the current real-world scenario. Of course, in addition to the few-shot object detection model, other object detection models can also be trained, such as the YOLO (You Only Look Once) model and the SSD (Single Shot MultiBox Detector) model.

[0111] For the above-mentioned target detection and annotation function, it is necessary to provide a method to improve the efficiency of annotating training images, thereby improving the efficiency of training target detection models. Therefore, an embodiment of the present application provides an image annotation method, a model training method, a device, and an electronic device. The following first introduces the image annotation method provided by the embodiment of the present application. The method can be applied to an electronic device, and the electronic device can be integrated with a display screen to display the interactive interface in this embodiment. For example, the electronic device can be a computer, a tablet computer, a smart phone, etc.

[0112] See also Figure 2 , Figure 2 To utilize Figure 1 The schematic diagram of the principle flow of training the target detection model using the few-shot target detection training operator shown in FIG. Figure 2 As shown, this process may include the following steps:

[0113] S201, add 1-10 training images with significant feature differences;

[0114] S202: Create a corresponding number of target type labels according to the actual scenario requirements;

[0115] S203, creating a target sample for each type of target;

[0116] S204, marking all targets in the remaining images through "intelligent append" or "automatic marking";

[0117] S205, training the images added to the target training set (based on the deep learning model).

[0118] In the process of deep learning model training, in order to improve the annotation accuracy and include enough actual scenes, hundreds of sample images are needed as samples for training, which is time-consuming and laborious. Although ordinary few-sample learning training reduces the number of samples needed, repeated annotation operations are still unavoidable. The embodiment of the present application combines the three methods of few-sample detection training, intelligent appending, and automatic annotation to annotate multiple images and multiple target objects, effectively improving the work efficiency of target detection and annotation in complex scenes.

[0119] based on Figure 2 According to the principle shown, the embodiment of the present invention provides two specific labeling methods: the first method is to label the training images manually in combination with automatic labeling and other functions; the second method is to label the training images manually in combination with intelligent appending functions.

[0120] The following sections provide detailed explanations of each.

[0121] First, the first method of labeling training images is described in detail.

[0122] See also Figure 3 , Figure 3 This is a flow chart of a first embodiment of the first image annotation method provided in the embodiment of the present application, as shown in FIG. Figure 3 As shown, the image annotation method provided in the embodiment of the present application includes the following steps:

[0123] S301, displaying a first training image to be labeled;

[0124] S302, receiving user annotations for a first training image, and displaying target boxes and category labels of different target objects annotated by the user on the first training image; wherein each category label annotates at least one target object;

[0125] S303, in response to the automatic labeling instruction, displaying on the second training image: a target box and a category label of the target object in the second training image; wherein the target box and the category label of the target object in the second training image are obtained by inputting the second training image into a target detection model trained using the labeled first training image; the labeled first training image includes: a target box and a category label labeled by the user.

[0126] In this embodiment, the user can first manually annotate the first training image displayed by the electronic device, allowing the electronic device to train the target detection model using the annotated first training image. In this way, the user can then use automatic labeling instructions to instruct the electronic device to input the second training image into the trained target detection model, thereby automatically obtaining the target box and category label of the target object in the second training image. It can be seen that, using this embodiment of the present invention, only the first training image needs to be manually annotated, and the second training image can be automatically annotated by the electronic device. Therefore, this solution can more efficiently annotate images and reduce the time cost of annotation.

[0127] In some embodiments, if the training atlas contains both the first training image and the second training image, before automatically labeling the second training image, the user can also intelligently append the first training image to display more target boxes and category labels on the first training image, and then automatically label the second training image.

[0128] For details, see Figure 4 , Figure 4 This is a flow chart of the second embodiment of the first image annotation method provided in the embodiment of the present application, as shown in FIG. Figure 4 As shown, the method includes:

[0129] S401, obtaining a training atlas determined by a user; wherein the training atlas includes a plurality of training images;

[0130] In this embodiment, the electronic device can display an interactive interface for annotating training images. Referring to FIG5(a), FIG5(a) is a first schematic diagram of the interactive interface used by the image annotation method provided in the embodiment of the present application (in the state where no image is imported). As shown in FIG5(a), the user can, in the "1 / Add Training Image" area, click "Camera Capture" to control the camera connected to the electronic device to capture the training image in real time; can also click "Stored Image Import" to select an image from images pre-stored in the electronic device as a training image; can also click "External Import" to obtain a training image from an external device connected to the electronic device.

[0131] While displaying the interactive interface shown in Figure 5 (a), the electronic device can also display a dialog box for setting commonly used parameters for target detection, see Figure 5 (b). Figure 5 (b) is a second schematic diagram of the interactive interface used by the image annotation method provided in the embodiment of the present application (displaying the status of commonly used parameters for target detection). As shown in Figure 5 (b), the dialog box includes: a switch control corresponding to angle enable. The user can click the switch control to turn on the angle enable or turn off the angle enable. When the angle enable is turned on, a rectangular box with an angle value will be used to mark the target object in the training image of the target detection model used for training; when it is turned off, a rectangular box without an angle will be used to mark the target object. It can be understood that a rectangular box without an angle refers to a horizontal and vertical rectangular box.

[0132] The dialog box may also include: a switch control for enabling the optimal model resolution setting. This parameter may not be enabled by default. If the user thinks that the display effect of the current training image is not good, the user can click the switch control to turn on the parameter. After turning it on, the "Optimal Model Resolution" parameter will be displayed in the dialog box. After that, the user can set the value of the optimal model resolution in the dialog box. In this way, when the training image is displayed in the subsequent interactive interface, the number of pixels on the longer side of the displayed training image will be adjusted according to the value set by the user, and the other side will be scaled proportionally according to the set value. For example, if the setting is 512, the number of pixels on the longer side of the displayed training image will be 512.

[0133] When the target detection model to be trained is a few-sample target detection model, if the first method of annotating training images is used, the user can select 2-10 training images as the training atlas. Of course, more than 10 training images are also possible. The user can try to select images with significant feature differences to increase the diversity of the samples and thus improve the training effect of the target detection model. Each training image can contain only one category of target objects or multiple categories. If the second method of annotating training images is used, the user can select 1-10 training images as the training atlas. The specific process will be described in detail below.

[0134] S402, receiving a training image selected by a user from a training atlas as a first training image to be labeled; wherein, except for the first training image to be labeled, the other images in the training atlas are used as second training images to be labeled;

[0135] After the user determines the training atlas, they can select an image that includes all target objects of all categories as the first training image. By obtaining the user-determined training atlas and receiving the user-selected training image from the training atlas as the first training image to be annotated, a foundation can be laid for subsequent annotation of the first and second training images.

[0136] S403, displaying a first training image to be labeled;

[0137] In this step, the user can select an image as the first training image after obtaining the training atlas determined by the user. After the user selects the first training image, the first training image can be displayed in the user interaction interface of the electronic device. For other training images that are not currently displayed, the electronic device can display their thumbnails in the thumbnail display area. Figure 6 , Figure 6 This is the third schematic diagram of the interactive interface shown in Figure 5 (a) (the state of importing images). Figure 6 As shown, the thumbnail corresponding to the currently displayed training image can display a stroke effect to prompt the user which training image is currently displayed.

[0138] S404, in response to the label creation instruction, receiving and saving different category labels corresponding to different target objects created by the user for the first training image;

[0139] Before responding to the label creation instruction, the electronic device may display a label creation control; thus, the label creation instruction may be generated by the user clicking the label creation control. Figure 6 For example, the label creation control can be "New" in the "Label List" area. After that, the electronic device can display the label creation dialog box, see Figure 7 , Figure 7 This is the fourth schematic diagram of the interactive interface shown in Figure 5 (a) (the state when creating a tag), as shown in Figure 7 As shown, users can enter the label name in this dialog box and set the label color, so that different category labels can be displayed in different colors, making it easier for users to distinguish.

[0140] For example, if the first training image contains two different target objects: a large pill and a small pill, the user can create two category labels and enter the label names as "large pill" and "small pill" respectively. After the user completes creating the category labels, the label controls corresponding to the created category labels can be displayed under "Label Category". Figure 8 , Figure 8 This is the fifth schematic diagram of the interactive interface shown in Figure 5 (a) (the state after the label is created). Figure 8 As shown, the label control corresponding to each category label may further include "Edit" and "Delete" buttons, which are used to modify the name of the category label and delete the category label respectively.

[0141] This embodiment receives instructions by displaying controls, which allows users to trigger corresponding functions more conveniently and improves user experience.

[0142] S405, receiving user annotations for the first training image, and displaying target boxes and category labels of different target objects annotated by the user on the first training image; wherein each category label annotates at least one target object;

[0143] by Figure 8 For example, the user can click "Rectangular Box Selection" in the "2 / Annotate Image" area to annotate the target object in the first training image. For example, the user can draw a rectangle as a target box at the location of the target object, and enclose the target object inside the rectangle. The electronic device displays the target box drawn by the user, and displays the label category of the target object near the target box.

[0144] Before receiving the user's annotations for the first training image, the electronic device may first display label controls corresponding to different target objects created by the user. In this case, the target box and category label for each target object annotated by the user on the first training image are obtained by selecting the label control corresponding to the target object.

[0145] When there are multiple target objects in the first training image, the user can first select the label control corresponding to the created target object. The selected label control can also be displayed in the form of a stroke, and then the user can label the target object in the training image. Afterwards, the user can also select the label control corresponding to other created target objects to label target objects of other categories, thereby achieving labeling of target objects of different categories. Figure 9 , Figure 9 This is the sixth schematic diagram of the interactive interface shown in FIG5(a) (the state after the user has annotated the first training image). Figure 9As shown, the user can label a "large pill" and a "small pill" in the image. The target boxes and category labels marked by the user can be used as target examples. Users only need to label one target object per category. Users can also label multiple target objects per category to effectively enrich the target examples and make subsequent annotations more accurate.

[0146] S406 , in response to the label appending instruction, displaying on the first training image: a target frame and a category label of the target object in the first training image;

[0147] The target bounding box and category label of the target object in the first training image are obtained by inputting the first training image into the target detection model trained using the first training image annotated by the user;

[0148] Before responding to the tag append instruction, the electronic device may first display a smart append control; in this way, the tag append instruction can be generated by the user clicking the smart append control. After the user clicks the smart append control, responding to the tag append instruction may take some time. During this time, the electronic device may provide a loading wait prompt, such as displaying a dialog box containing text to prompt the user.

[0149] In one implementation, responding to the label appending instruction includes: responding to the label appending instruction when the smart appending control is unlocked; wherein the unlocking condition of the smart appending control is: the presence of a user-annotated target box and category label on the first training image.

[0150] That is to say, when there are no user-annotated target boxes and category labels in the training atlas, the smart additional control can be in a locked state. For example, the electronic device can temporarily not display the smart additional control, or display the smart additional control in a lighter color until the smart additional control is unlocked, and then display the smart additional control normally. By locking and unlocking the control, the user can avoid accidentally touching the control when the function corresponding to the control cannot take effect.

[0151] If unlabeled target objects still exist in the user-labeled first training image, the user can issue a label-adding instruction. Upon receiving the label-adding instruction, the electronic device can first train the target detection model based on the target boxes and category labels annotated by the user. The first training image is then input into the target detection model to obtain and display the target boxes and category labels of the target objects detected by the target detection model in the first training image.

[0152] The target detection model can be trained by few-shot detection training. For example, the training process of the target detection model can be as follows:

[0153] First, an unlabeled first training image is input into the target detection model to be trained, and the target box and category label of the target object detected by the model in the image are obtained as the model prediction result; the model loss is calculated based on the model prediction result and the true value of the first training image, and the model parameters of the target detection model are adjusted using the calculated model loss until the model converges, thereby obtaining the trained target detection model; wherein the true value of the first training image can be the target box and category label that have been labeled in the first training image.

[0154] S407, displaying the second training image to be labeled;

[0155] After completing the annotation of the first training image, the user can switch to the second training image, so that the electronic device displays the second training image. Figure 6 As shown, the user can click on the thumbnail of the training image that is not currently displayed in the thumbnail display area to switch the training image.

[0156] S408 , in response to the automatic labeling instruction, displaying on the second training image: a target frame and a category label of the target object in the second training image;

[0157] The target bounding boxes and category labels of the target objects in the second training image are obtained by inputting the second training image into an object detection model trained using the labeled first training image. The labeled first training image includes the user-annotated target bounding boxes and category labels. Of course, if the electronic device responds to a label appending instruction for the first training image, the labeled first training image also includes the appended target bounding boxes and category labels.

[0158] Similarly, before responding to the automatic marking instruction, the electronic device may first display the automatic marking control, so that the automatic marking instruction can be generated by the user by clicking the automatic marking control.

[0159] The step of responding to the automatic marking instruction may include: responding to the automatic marking instruction when the automatic marking control is unlocked; wherein the unlocking condition of the automatic marking control is: the first training image has been marked and the second training image to be marked is displayed.

[0160] In one implementation, see Figure 10 ,Figure 5(a) shows the seventh schematic diagram of the interactive interface (showing the ,status of the automatic marking button), as shown in Figure 10As shown, while displaying the second training image, a toast (a non-modal pop-up window) is generated. This prompt displays the automatic labeling control ("Train and Automatically Label" in the image) and the text "Automatically label images using automated reasoning." The specific presentation of the automatic labeling control can be customized based on your business scenario; this is just an example.

[0161] After receiving the automatic labeling instruction, the electronic device can first train the object detection model using the labeled first training image. Then, the second training image is input into the trained object detection model to obtain and display the target bounding box and category label of the target object in the second training image. Because this process is time-consuming, the electronic device may also provide a loading waiting prompt during this period.

[0162] In a specific implementation, step S408 may include:

[0163] In response to the automatic labeling instruction, the target boxes and category labels of the target objects marked after the second training image is input into the target detection model are displayed on the second training image; and based on the user's selection, the target boxes and category labels of all target objects corrected and adjusted by the user are displayed on the second training image.

[0164] Specifically, the correction and adjustment process may include the following situations:

[0165] Case 1: When a modification instruction for any target box is received, the position, size, and / or corresponding category label of the target box is modified based on the modification instruction;

[0166] Case 2: When a deletion instruction is received for any target box, the target box and the corresponding category label are removed from the training image;

[0167] Case three: when a new annotation added by the user for a training image is received, the target box and category label added by the user are displayed on the training image.

[0168] In this embodiment, the electronic device supports the user to modify and adjust the target box and category label of the target object, which can further ensure the accuracy of the labeling and thereby improve the training effect of the target detection model.

[0169] When the number of the second training images to be labeled is greater than one, after executing the automatic labeling instruction on the first second training image, the method may further include:

[0170] S409 , sequentially displaying the next second training image to be labeled, and executing the automatic labeling instruction for the second training image until all second training images are labeled.

[0171] In other words, steps S407-S408 can be performed multiple times until all second training images are labeled. This allows for rapid labeling of multiple training images. This solution provides an automatic labeling function for multi-image scenarios, enabling rapid identification of objects of the same type across multiple images. This effectively reduces labeling time and improves efficiency in complex scenarios with numerous duplicate objects.

[0172] In addition, this embodiment combines three methods: few-sample detection training based on deep learning samples, intelligent appending, and automatic labeling. This reduces the number of manual labeling required from hundreds or thousands to dozens, greatly reducing the time and energy consumed by manual labeling during target detection training, and also greatly reducing the time and labor costs of actual application deployment.

[0173] In this embodiment, the user can create a label in the electronic device, and then manually label the first training image displayed by the electronic device, so that the electronic device can use the labeled first training image to train the target detection model. In this way, the electronic device can automatically label the target objects in the first training image that the user has not labeled, and after displaying the second training image, the user can use the automatic labeling instruction to make the electronic device input the second training image into the trained target detection model, thereby automatically obtaining the target box and category label of the target object in the second training image. It can be seen that through this solution, only part of the target objects in the first training image need to be manually labeled, and the remaining target objects in the first training image and the target objects in the second training image can be automatically labeled by the electronic device. Therefore, through this solution, images can be labeled more efficiently and the time cost of labeling can be reduced.

[0174] In one embodiment of the present application, see Figure 11 , Figure 11 for Figure 4 The process diagram of intelligent addition in the embodiment shown is as follows: Figure 11 As shown, the process may include the following steps:

[0175] S1101, in response to a label appending instruction, displays a target frame and a category label of the target object additionally annotated this time on the first training image; wherein, the target frame and the category label of the target object additionally annotated this time are obtained by inputting the first training image into a target detection model trained using the first training image annotated by the user, and the similarity between the target object additionally annotated this time and the target object annotated by the user is greater than a preset threshold.

[0176] Among them, the similarity between two target objects may refer to: the image similarity between the local images framed by the target boxes that mark the two target objects. That is to say, after the electronic device responds to the label appending instruction, it may not append the target boxes and category labels of all target objects detected by the target detection model, but only append the target objects whose similarity with the target objects marked by the user is greater than a preset threshold. Specifically, the electronic device may calculate the similarity between each target object under each category detected by the target detection model and the target objects of the category that are currently marked. If there are multiple target objects of the category that are currently marked, the average value of the similarities may be taken. If the similarity is greater than the preset threshold, the target box and category label of the target object will be appended, otherwise no additional labeling will be added.

[0177] The preset threshold can be set according to actual conditions and experience. In one implementation, the preset threshold can be set higher, such as 98%, so that only 1-3 target objects are automatically added and marked each time. This number is convenient for users to confirm and adjust and gradually release a small number of inference targets, which helps to continuously improve the accuracy of reasoning.

[0178] For example, see FIG12(a), which is a schematic diagram of label addition in the image annotation method provided in the embodiment of the present application. Figure 9 After performing the first additional annotation on the first training image in , the obtained first training image can be shown as FIG12( a ).

[0179] After each additional annotation, the electronic device may further modify and adjust the target frame and category label of the target object of this additional annotation based on user selection.

[0180] After the label appending instruction is executed, step S1102 is executed: determining whether the target frames and category labels of all target objects are displayed on the first training image; if not, executing S1103; if so, executing S1104.

[0181] When the first training image does not display target frames and category labels of all target objects, the method may further include:

[0182] S1103, in response to the next label addition instruction, displaying a target bounding box and a category label of the newly annotated target object on the first training image; wherein the target bounding box and the category label of the newly annotated target object are obtained by inputting the first training image into the updated target detection model, and the similarity between the newly annotated target object and the previously annotated target object is greater than a preset threshold;

[0183] The updated target detection model is obtained by training the target detection model based on the first training image containing the target box and category label annotated by the user, as well as the target box and category label annotated previously.

[0184] That is, the user can issue label appending instructions multiple times until all target objects in the first training image are labeled, or until the target detection model cannot detect any new target objects. Furthermore, before each response to the label appending instruction, the electronic device can retrain the target detection model using the target boxes and category labels labeled by the user, as well as the first training image with the previously labeled target boxes and category labels, to obtain an updated target detection model.

[0185] For example, if there are 10 target objects in the first training image, the user can manually mark one target object in the first training image and then issue a label-adding instruction. In response to the label-adding instruction, the electronic device additionally marks three target objects. The user can then issue another label-adding instruction. In response to the label-adding instruction, the electronic device further marks two target objects, and so on. This cycle continues until the electronic device has finally marked all target objects, at which point the user no longer needs to issue a label-adding instruction.

[0186] For example, see FIG12( b ), which is a schematic diagram showing that all target objects are marked in the image marking method provided in an embodiment of the present application.

[0187] S1104, completing the intelligent appending process.

[0188] In this embodiment, the electronic device can gradually release the results of the additional annotation, reserving enough for the user to gradually correct the target detection model, which can greatly improve the accuracy of the additional annotation process and reduce detection errors caused by external factors such as the environment.

[0189] For easier understanding, see Figure 13 , Figure 13 for Figure 4 The schematic diagram of the intelligent additional interaction principle of the embodiment shown is as follows: Figure 13 As shown, the process includes the following steps:

[0190] S1301, the smart add function is disabled by default; that is, in the interactive interface initially displayed by the electronic device, the smart add control is not unlocked.

[0191] S1302: At least one target has been marked in the current image, and the unlocking condition of the smart additional function has been met, then S1303 is executed;

[0192] S1303, the smart add function is released from the disabled state;

[0193] S1304, click "Smart Add"; that is, the user clicks "Smart Add".

[0194] S1305: Call the algorithm to mark the most similar targets (i.e., target objects) located on the image based on the existing annotation content. The number is 1-3.

[0195] This step is the aforementioned process of using the labeled first training image to train the target detection model and then inputting the first training image into the target detection model.

[0196] S1306, determining whether the added content is correct, that is, the user determines whether the target object added intelligently is correct; if so, executing S1307; if not, executing S1308;

[0197] S1307, determining whether all objects in the current image have been added; that is, the user determines whether all objects in the current image have been added; if so, executing S1309; if not, returning to executing 1304;

[0198] S1308 , modifying and adjusting the marked content; that is, the user modifies and adjusts the target frame and category label of the marked target object.

[0199] S1309, switch to the next image.

[0200] In one embodiment of the present application, see Figure 14 , Figure 14 for Figure 4 The process diagram of automatic marking in the embodiment shown is as follows: Figure 14 As shown, the process may include the following steps:

[0201] S1401, displaying a first second training image to be labeled;

[0202] S1402, in response to the automatic labeling instruction, displaying on the second training image the target boxes and category labels of the target objects marked after the second training image is input into the target detection model; and, based on the user's selection, displaying on the second training image the target boxes and category labels of all target objects corrected and adjusted by the user.

[0203] After executing the automatic labeling instruction on each second training image, the image labeling method provided in the embodiment of the present application may further include:

[0204] S1403, based on the user's selection, adding the second training image that has been corrected and adjusted by the user to the target training set to update the target training set;

[0205] Specifically, after the user determines that a second training image has been labeled, the user can decide whether to add the second training image to the target training set. If the user determines that the second training image needs to be added to the target training set, the electronic device can first determine whether the second training image has been manually corrected and adjusted by the user. If the second training image has not been corrected and adjusted, the electronic device can refuse to add the second training image to the target training set and display a text prompt to the user. For example, see Figure 15 , Figure 15 This is the eighth schematic diagram of the interactive interface shown in Figure 5 (a) (the state of prompting the rejection of the image to be added to the target training set), as shown in Figure 15 As shown, text can be displayed below the second training image: "Images that have been automatically labeled will not participate in training and will be added to training samples after manual adjustment."

[0206] If the second training image is corrected and adjusted, the second training image can be added to the target training set.

[0207] S1404, displaying the next second training image to be labeled;

[0208] After displaying the next second training image to be labeled each time and before responding to the automatic labeling instruction, the image labeling method may further include:

[0209] S1405, based on the user selection, training the target detection model with the updated target training set to update the target detection model;

[0210] That is to say, after each display of the next second training image, the user can decide whether to update the target detection model. If updating is required, the electronic device can use all images in the current target training set to update the target detection model; if updating is not required, the target detection model that was automatically labeled for the previous second training image can continue to be used.

[0211] For the second and subsequent second training images, in response to the automatic labeling instruction, the target frame and category label of the target object in the second training image are displayed on the second training image, including:

[0212] S1406, based on user selection, displaying on the second training image the target boxes and category labels of the target objects marked after the second training image is input into the unupdated or updated target detection model; and, based on user selection, displaying on the second training image the target boxes and category labels of all target objects corrected and adjusted by the user.

[0213] See also Figure 16 , Figure 16 This is the ninth schematic diagram of the interactive interface shown in Figure 5 (a). Figure 16 As shown, after each display of the next second training image, the electronic device can display a toast prompt box with two controls: "Auto Label" and "Retrain and Auto Label." Of course, this process is not limited to this display method, and this is only an example. If the user clicks "Auto Label," it means that the target detection model does not need to be updated, and the second training image inputs the unupdated target detection model; if the user clicks "Retrain and Auto Label," it means that the target detection model needs to be updated, and the second training image inputs the updated target detection model.

[0214] For easier understanding, see Figure 17 , Figure 17 for Figure 4 The schematic diagram of the automatic marking interaction principle of the embodiment shown is as follows: Figure 17 As shown, the process may include the following steps:

[0215] S1701, the automatic marking function is disabled by default; that is, the automatic marking function is not unlocked in the interactive interface initially displayed by the electronic device;

[0216] S1702, content has been annotated on an image;

[0217] When switching to an image without a label, the automatic labeling function unlocking condition is met, and S1703 is executed;

[0218] S1703, the automatic marking function is released from the disabled state, and the relevant instructions are displayed on the interactive interface; For relevant instructions, see Figure 10 .

[0219] S1704, click "Train and Automatically Label". That is, the user clicks "Train and Automatically Label".

[0220] S1705, marking all inferred targets on the image; that is, the electronic device will train the model based on the existing annotation content, and then mark all inferred target objects on the image.

[0221] The first time this step is executed, it is the process of training the object detection model using the annotated first training image and then inputting the second training image into the object detection model. For the second and subsequent executions, this step is the process of displaying the target bounding boxes and category labels of the target objects annotated on the second training image, either after inputting the unupdated or updated object detection model, based on the user's selection.

[0222] S1706, determining whether to adjust the automatic annotation; that is, the user can determine whether to adjust the automatic annotation; if so, executing S1707; if not, executing S1708.

[0223] S1707 , adjusting the marked content or supplementing the targets that are not automatically marked; that is, the user adjusts the target frame and category label of the marked target object, or supplements the target object that is not automatically marked.

[0224] S1708, confirm whether to include the current image in the target training set; that is, the user confirms whether to include the current image in the target training set; if so, execute S1707, at this time the interactive interface will display a prompt message if the image has not been manually adjusted, to remind the user that the image included in the target training set needs to be manually adjusted. If not, execute S1709;

[0225] S1709, switching to the next image without a label;

[0226] S1710, confirm whether to use the current model for automatic marking; that is, when switching to the next image, the interactive interface will display a prompt message to allow the user to confirm whether to use the current model for automatic marking, such as Figure 16 If yes, execute S1711; if no, execute S1712;

[0227] S1711, continue automatic labeling with the current model; at this time, the user can execute S1713: click "Automatic Labeling", and then the interaction logic will return to S1705 again;

[0228] S1712: Retrain a new model based on the current target training set and then perform automatic labeling. This means updating the target detection model and using it for automatic labeling. At this point, the user can proceed to S1714: Click "Retrain and Automatically Label." The interactive logic will then return to S1705. This cycle will complete labeling for all images.

[0229] In this embodiment, the user can screen the training images to determine whether to add the training images to the target training set. The user can then use the images in the target training set to update the target detection model to continuously improve the performance of the target detection model, making the automatic labeling function more and more accurate.

[0230] The second method for labeling training images is described in detail below.

[0231] The present application also provides an image annotation method, which is applied to electronic devices. Figure 18 , Figure 18This is a flow chart of the first embodiment of the second image annotation method provided in the embodiment of the present application, as shown in FIG. Figure 18 As shown, the method includes:

[0232] S1801, displaying a training image to be labeled;

[0233] S1802, receiving user annotations for the training image, and displaying target boxes and category labels of different target objects annotated by the user on the training image;

[0234] Among them, each category label is annotated with at least one target object;

[0235] S1803 , in response to the label appending instruction, displaying on the training image: a target frame and a category label of the target object in the training image;

[0236] The target frame and category label of the target object in the training image are obtained by inputting the training image into a target detection model trained using the training image annotated by the user.

[0237] The above steps are similar to S403-S406, and the similarities are not repeated here. Figure 4 The difference between the illustrated embodiment is that in this embodiment, the image annotation can be completed by using the intelligent appending function without using the automatic annotation function.

[0238] In this embodiment, only some target objects in the training image need to be manually labeled, and other target objects can be automatically labeled by the electronic device. Therefore, this solution can more efficiently label the image and reduce the time cost of labeling.

[0239] See also Figure 19 , Figure 19 This is a flow chart of a second embodiment of the second image annotation method provided in the embodiment of the present application, as shown in FIG. Figure 19 As shown, the method includes:

[0240] S1901, displaying a training image to be labeled;

[0241] S1902, receiving annotations for the training image from a user, and displaying target boxes and category labels of different target objects annotated by the user on the training image; wherein each category label annotates at least one target object;

[0242] S1903, in response to the label appending instruction, displaying on the training image: a target frame and a category label of the target object in the training image;

[0243] The target box and category label of the target object in the training image are obtained by inputting the training image into the target detection model trained using the training image annotated by the user;

[0244] The above steps are similar to S403-S406, and the similarities are not repeated here.

[0245] After step S1903 , the above steps S1102 - S1103 may be further performed until the target boxes and category labels of all target objects are displayed on the current training image, and then step S1904 is performed.

[0246] S1904, determine whether there are any training images to be labeled; if not, execute S1906; if so, execute S1905;

[0247] S1905: Display the next training image to be labeled; then return to step S1902.

[0248] S1906, completing the intelligent append process.

[0249] In other words, if there are multiple training images to be labeled, after executing the label appending command for the first training image, the next training image to be labeled can be displayed in sequence. After the next training image to be labeled is displayed, it can also be manually labeled, and then the label appending command can be executed for this training image again. This cycle continues until all training images are labeled.

[0250] In this embodiment, only some target objects in the training image need to be manually labeled, and other target objects in the training image can be automatically labeled by the electronic device. Therefore, this solution can more efficiently label the image and reduce the time cost of labeling.

[0251] In one embodiment, in the second image annotation method provided in the embodiment of the present application, the user can also annotate the training image by automatic annotation. Figure 20 , Figure 20 This is a flow chart of a third embodiment of the second image annotation method provided in the embodiment of the present application, as shown in FIG. Figure 20 As shown, the method includes:

[0252] S2001, displaying a training image to be labeled;

[0253] S2002, receiving annotations for the training image from a user, and displaying target frames and category labels of different target objects annotated by the user on the training image; wherein each category label annotates at least one target object;

[0254] S2003, in response to the label appending instruction, displaying on the training image: a target frame and a category label of the target object in the training image;

[0255] The target box and category label of the target object in the training image are obtained by inputting the training image into the target detection model trained using the training image annotated by the user;

[0256] The above steps are similar to S403-S406, and the similarities are not repeated here.

[0257] After step S2003 , the above steps S1102 - S1103 may also be performed until the target frames and category labels of all target objects are displayed on the current training image, and then step S2004 is performed.

[0258] S2004, determine whether there are any training images to be labeled; if not, execute S2007; if so, execute S2005;

[0259] S2005, displaying the next training image to be labeled;

[0260] S2006, in response to the automatic labeling instruction, displays on the next training image: the target box and category label of the target object in the training image; wherein, the target box and category label of the target object in the training image are obtained by inputting the training image into the target detection model trained using the labeled training image; then the process can return to step S2004 until all training images are labeled.

[0261] S2007, complete the automatic labeling process.

[0262] Based on the above image annotation method, the present application embodiment also provides a model training method, see Figure 21 , Figure 21 This is a flow chart of the second embodiment of the model training method provided in the embodiment of the present application, as shown in FIG. Figure 21 As shown, the method includes the following steps:

[0263] S2101, obtaining a training image determined by a user;

[0264] S2102, applying any of the above-mentioned image annotation methods to annotate the obtained training image;

[0265] S2103 , in response to the model training instruction, the target detection model is trained using the labeled training images to obtain a trained target detection model.

[0266] The model training instruction can also be issued by the user by clicking the corresponding control in the interactive interface. For example, the control can be Figure 6 , click the Train button.

[0267] The present application embodiment does not limit the specific training method. For example, the training can be performed in the following manner:

[0268] First, the unlabeled training image is input into the target detection model to be trained, and the target box and category label of the target object detected by the model for the image are obtained as the model prediction result; the model loss is calculated based on the model prediction result and the labeling result of the training image, and the calculated model loss is used to adjust the model parameters of the target detection model until the model converges, and a trained target detection model is obtained.

[0269] In this embodiment, only some target objects in the training image need to be manually labeled, and the remaining target objects in the training image and the target objects in other training images can be automatically labeled by the electronic device. Therefore, this solution can more efficiently label the image, reduce the time cost of labeling, and thus improve the efficiency of target detection model training.

[0270] The present application also provides an image annotation device, which is applied to electronic devices. Figure 22 , Figure 22 This is a structural diagram of the first embodiment of the image annotation device provided in the embodiment of the present application; Figure 22 As shown, the device includes:

[0271] Display module 2201, configured to display a first training image to be labeled;

[0272] The annotation receiving module 2202 is configured to receive user annotations for the first training image and display target boxes and category labels of different target objects annotated by the user on the first training image; wherein each category label annotates at least one target object;

[0273] The automatic labeling response module 2203 is used to display on the second training image in response to the automatic labeling instruction: the target box and category label of the target object in the second training image; wherein the target box and category label of the target object in the second training image are obtained by inputting the second training image into the target detection model trained using the labeled first training image; the labeled first training image includes: the target box and category label labeled by the user.

[0274] In one embodiment, the device further includes: a label appending response module, which is used to receive the user's annotations for the first training image in the annotation receiving module 2202, and after displaying the target frames and category labels of different target objects annotated by the user on the first training image, in response to the label appending instruction, display on the first training image: the target frames and category labels of the target objects in the first training image; wherein the target frames and category labels of the target objects in the first training image are obtained by inputting the first training image into a target detection model trained using the first training image annotated by the user; the annotated first training image also includes: the target frames and category labels annotated after the appending.

[0275] In one embodiment, the apparatus further comprises:

[0276] The label creation response module is used to respond to the label creation instruction before the label receiving module 2202 receives the user's label for the first training image, and receive and save different category labels corresponding to different target objects created by the user for the first training image.

[0277] In one embodiment, the device further includes: an image switching module, which is used to, when the number of second training images to be labeled is greater than 1, display the next second training image to be labeled in sequence after the automatic labeling response module 2203 executes the automatic labeling instruction on the first second training image, and execute the automatic labeling instruction on the second training image until all second training images are labeled.

[0278] In one embodiment, the apparatus further comprises: a training atlas acquisition module, configured to obtain a training atlas determined by a user before the display module 2201 displays the first training image to be annotated; wherein the training atlas includes a plurality of training images;

[0279] The training image receiving module is used to receive a training image selected by the user from the training atlas as the first training image to be labeled; wherein, except the first training image to be labeled, the other images in the training atlas are used as the second training images to be labeled.

[0280] In one embodiment, the automatic marking response module 2203 is specifically configured to:

[0281] For the first second training image, in response to the automatic labeling instruction, the target box and category label of the target object marked after the second training image is input into the target detection model are displayed on the second training image; and based on the user's selection, the target box and category label of all target objects corrected and adjusted by the user are displayed on the second training image.

[0282] In one embodiment, the apparatus further comprises: a target training set updating module configured to, after executing the automatic marking instruction on a second training image each time, add the second training image corrected and adjusted by the user to the target training set based on a user selection, so as to update the target training set;

[0283] a model updating module configured to train the target detection model with the updated target training set based on user selection to update the target detection model after the image switching module displays the next second training image to be labeled each time and before the automatic labeling response module 2203 responds to the automatic labeling instruction;

[0284] The automatic marking response module 2203 is specifically used to:

[0285] Based on user selection, the target boxes and category labels of the target objects marked after the second training image is input into the unupdated or updated target detection model are displayed on the second training image; and based on user selection, the target boxes and category labels of all target objects after user correction and adjustment are displayed on the second training image.

[0286] In one embodiment, the tag appending response module is specifically configured to:

[0287] In response to the label appending instruction, the target frame and category label of the target object additionally annotated this time are displayed on the first training image; wherein, the target frame and category label of the target object additionally annotated this time are obtained by inputting the first training image into the target detection model trained using the first training image annotated by the user, and the similarity between the target object additionally annotated this time and the target object annotated by the user is greater than a preset threshold.

[0288] In one embodiment, the label appending response module is also used to, after the label appending instruction is executed, in the case that the target frames and category labels of all target objects are not displayed on the first training image, in response to the next label appending instruction, display the target frames and category labels of the target objects additionally annotated on the first training image; wherein the target frames and category labels of the target objects additionally annotated are obtained by inputting the first training image into the updated target detection model, and the similarity between the target objects additionally annotated and the already annotated target objects is greater than a preset threshold; the updated target detection model is obtained by training the target detection model based on the first training image containing the target frames and category labels annotated by the user, and the target frames and category labels previously annotated.

[0289] In one embodiment, the apparatus further comprises:

[0290] The correction and adjustment module is used to correct and adjust the target frame and category label of the target object additionally labeled this time based on user selection after the label addition response module displays the target frame and category label of the target object additionally labeled this time.

[0291] In one embodiment, the apparatus further comprises:

[0292] A first control display module is configured to display a label creation control before the label creation response module responds to a label creation instruction; wherein the label creation instruction is generated by the user by clicking the label creation control;

[0293] a second control display module configured to display label controls corresponding to different target objects created by the user before the label receiving module 2202 receives the user's annotations for the first training image; wherein the target box and category label of each target object annotated by the user on the first training image are obtained by annotating the target object when the user selects the label control corresponding to the target object;

[0294] A third control display module is configured to display an automatic marking control before the automatic marking response module 2203 responds to an automatic marking instruction; wherein the automatic marking instruction is generated by the user by clicking on the automatic marking control;

[0295] The fourth control display module is used to display a smart append control before the label append response module responds to the label append instruction; wherein the label append instruction is generated by the user by clicking the smart append control.

[0296] In one embodiment, the automatic marking response module 2203 is specifically configured to: respond to an automatic marking instruction when the automatic marking control is unlocked; wherein the unlocking condition of the automatic marking control is: the first training image has been marked and the second training image to be marked is displayed;

[0297] The label appending response module is specifically used to respond to a label appending instruction when the smart appending control is unlocked; wherein, the unlocking condition of the smart appending control is that a target box and category label marked by the user exist on the first training image.

[0298] In one embodiment, the apparatus further comprises:

[0299] a modification response module, configured to, upon receiving a modification instruction for any target frame, modify the position, size, and / or corresponding category label of the target frame based on the modification instruction;

[0300] A deletion response module is configured to remove a target box and its corresponding category label from the training image upon receiving a deletion instruction for the target box;

[0301] The annotation adding module is used to display the target box and category label added by the user on the training image when receiving the annotation added by the user for the training image.

[0302] The present application also provides an image annotation device, which is applied to electronic devices. Figure 23 , Figure 23 This is a schematic structural diagram of a second embodiment of the image annotation device provided in an embodiment of the present application; the device includes:

[0303] The training image display module 2301 is used to display a training image to be labeled;

[0304] The user annotation receiving module 2302 is configured to receive user annotations for the training image and display target boxes and category labels of different target objects annotated by the user on the training image; wherein each category label annotates at least one target object;

[0305] The label appending instruction response module 2303 is used to respond to the label appending instruction and display on the training image: the target box and category label of the target object in the training image; wherein, the target box and category label of the target object in the training image are obtained by inputting the training image into a target detection model trained using the training image annotated by the user.

[0306] In one embodiment, the number of the training images to be labeled is multiple; the apparatus further includes:

[0307] The image appending module is used to display the next training image to be labeled in sequence after executing the label appending instruction for the first training image, and execute the label appending instruction for the training image until all training images are labeled.

[0308] In one embodiment, the number of the training images to be labeled is multiple; the apparatus further includes:

[0309] The automatic labeling module is used to display the next training image to be labeled in sequence after executing the label appending instruction for the first training image; in response to the automatic labeling instruction, display on the next training image: the target box and category label of the target object in the training image; wherein the target box and category label of the target object in the training image are obtained by inputting the training image into the target detection model trained using the labeled training image; and return to the step of sequentially displaying the next training image to be labeled until all training images are labeled.

[0310] The present application also provides a model training device for electronic equipment, see Figure 24 , Figure 24 This is a schematic diagram of the structure of the model training device provided in the embodiment of the present application, as shown in FIG. Figure 24 As shown, the device includes:

[0311] The training image obtaining module 2401 is used to obtain a training image determined by a user;

[0312] A training image annotation module 2402 is configured to apply any of the above-mentioned image annotation methods to annotate the obtained training images;

[0313] The model training instruction response module 2403 is used to respond to the model training instruction and use the labeled training images to train the target detection model to obtain a trained target detection model.

[0314] The present application also provides an electronic device. Figure 25 , Figure 25 This is a schematic diagram of the structure of the electronic device provided in the embodiment of the present application, such as Figure 25 Shown, including:

[0315] Memory 2501, used for storing computer programs;

[0316] The processor 2502 is configured to implement the steps of any of the above-mentioned image labeling methods or model training methods when executing the program stored in the memory 2501.

[0317] Furthermore, the electronic device may further include a communication bus and / or a communication interface, and the processor 2502, the communication interface, and the memory 2501 communicate with each other via the communication bus.

[0318] The communication bus mentioned in the electronic devices mentioned above can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus. This communication bus can be divided into address buses, data buses, control buses, etc. For ease of illustration, only a single thick line is used in the figure, but this does not mean that there is only one bus or only one type of bus.

[0319] The communication interface is used for communication between the above electronic device and other devices.

[0320] The memory may include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage. Alternatively, the memory may be at least one storage device located away from the processor.

[0321] The above-mentioned processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, and discrete hardware components.

[0322] In another embodiment provided in the present application, a computer-readable storage medium is also provided, which stores a computer program. When the computer program is executed by a processor, it implements the steps of any of the above-mentioned image annotation methods or model training methods.

[0323] In another embodiment provided by the present application, a computer program product comprising instructions is also provided, which, when executed on a computer, enables the computer to execute any image labeling method or model training method in the above embodiments.

[0324] In the above embodiments, all or part of the embodiments can be implemented using software, hardware, firmware, or any combination thereof. When implemented using software, all or part of the embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that integrates one or more available media. The available medium can be magnetic media (e.g., floppy disk, hard disk, tape), optical media (e.g., DVD), or solid-state drive (SSD).

[0325] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply the existence of any such actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or device comprising the element.

[0326] Each embodiment in this specification is described in a related manner. Similar portions between the various embodiments can be referenced to each other. Each embodiment focuses on the differences between the other embodiments. In particular, the device, electronic device, readable storage medium, and computer program product embodiments are generally similar to the method embodiments, so their descriptions are relatively simple. For related portions, reference can be made to the descriptions of the method embodiments.

[0327] The above description is only a preferred embodiment of the present application and is not intended to limit the scope of protection of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application are included in the scope of protection of the present application.

Claims

1. A method for labeling an image, characterized in that: Applied to electronic equipment, the method includes: Displaying an interactive interface for annotating training images; the interactive interface includes: an image display area, an option for adding training images, an option for annotating images, and a "training" button; After the user clicks the option to add a training image, the first training image to be labeled selected by the user is displayed in the image display area; After the user clicks the option to annotate the image, the user's annotations for the first training image are received, and target boxes and category labels of different target objects annotated by the user are displayed on the first training image; wherein each category label annotates at least one target object; Displaying a second training image to be labeled selected by a user in the image display area; An automatic labeling control is displayed on the interactive interface. In response to an automatic labeling instruction generated by a user clicking the automatic labeling control, a target frame and a category label of a target object in the second training image are displayed on the second training image. The target frame and category label of the target object in the second training image are obtained by inputting the second training image into an object detection model trained using the labeled first training image. The labeled first training image includes the target frame and category label labeled by the user. Based on the user selection, displaying the target boxes and category labels of all target objects corrected and adjusted by the user on the second training image; After the user clicks the "Train" button, the target detection model is trained using the labeled training images to obtain a trained target detection model.

2. The method according to claim 1, characterized in that After receiving the user's annotations for the first training image and displaying the target boxes and category labels of different target objects annotated by the user on the first training image, the method further includes: In response to a label appending instruction, displaying on the first training image: a target frame and a category label of a target object in the first training image; The target box and category label of the target object in the first training image are obtained by inputting the first training image into the target detection model trained using the first training image annotated by the user; The labeled first training image further includes: an additionally labeled target frame and category label.

3. The method according to claim 2, characterized in that Before receiving the user's annotation for the first training image, the method further includes: In response to the label creation instruction, different category labels corresponding to different target objects created by the user for the first training image are received and saved.

4. The method according to claim 1 or 2, characterized in that When the number of the second training images to be labeled is greater than one, after executing the automatic labeling instruction on the first second training image, the method further includes: The next second training image to be labeled is displayed in sequence, and the automatic labeling instruction is executed for the second training image until the labeling of all second training images is completed.

5. The method according to claim 4, characterized in that Before the step of displaying the first training image to be labeled, the method further includes: Obtaining a training atlas determined by a user; wherein the training atlas includes a plurality of training images; A training image selected by a user from the training atlas is received as a first training image to be labeled; wherein, except the first training image to be labeled, other images in the training atlas are used as second training images to be labeled.

6. The method according to claim 4, characterized in that For the first second training image, in response to the automatic labeling instruction, displaying on the second training image: a target frame and a category label of a target object in the second training image, including: In response to the automatic labeling instruction, the target box and category label of the target object marked after the second training image is input into the target detection model are displayed on the second training image.

7. The method according to claim 6, characterized in that After executing the automatic marking instruction on a second training image each time, the method further includes: adding the second training image that has been corrected and adjusted by the user to the target training set based on user selection to update the target training set; After displaying the next second training image to be labeled each time and before responding to the automatic labeling instruction, the method further includes: training the target detection model with the updated target training set based on the user selection to update the target detection model; For the second and subsequent second training images, in response to the automatic labeling instruction, displaying on the second training image: a target frame and a category label of the target object in the second training image, including: Based on the user's selection, displaying on the second training image the target box and category label of the target object annotated after the second training image is input into the unupdated or updated target detection model; and Based on the user selection, the target boxes and category labels of all target objects after user correction and adjustment are displayed on the second training image.

8. The method according to claim 2, characterized in that The step of displaying, in response to the label appending instruction, on the first training image: a target frame and a category label of a target object in the first training image, comprises: In response to the label appending instruction, the target frame and category label of the target object additionally annotated this time are displayed on the first training image; wherein, the target frame and category label of the target object additionally annotated this time are obtained by inputting the first training image into the target detection model trained using the first training image annotated by the user, and the similarity between the target object additionally annotated this time and the target object annotated by the user is greater than a preset threshold.

9. The method according to claim 8, characterized in that After the label appending instruction is executed, if target frames and category labels of all target objects are not displayed on the first training image, the method further includes: In response to the next label addition instruction, displaying a target bounding box and a category label of the newly annotated target object on the first training image; wherein the target bounding box and the category label of the newly annotated target object are obtained by inputting the first training image into the updated target detection model, and the similarity between the newly annotated target object and the previously annotated target object is greater than a preset threshold; The updated target detection model is obtained by training the target detection model based on the first training image including the target frame and category label annotated by the user and the target frame and category label annotated previously.

10. The method according to claim 9, characterized in that After displaying the target frame and category label of the target object that is additionally annotated, the method further includes: Based on the user's selection, the target box and category label of the target object to be annotated are corrected and adjusted.

11. The method according to claim 3, characterized in that Before responding to the label creation instruction, the method further includes: displaying a label creation control; wherein the label creation instruction is generated by the user by clicking the label creation control; Before receiving the user's annotations for the first training image, the method further includes: displaying label controls corresponding to different target objects created by the user; wherein the target box and category label of each target object annotated by the user on the first training image are obtained by the user selecting the label control corresponding to the target object; Before responding to the tag appending instruction, the method further includes: displaying a smart appending control; wherein the tag appending instruction is generated by the user by clicking the smart appending control.

12. The method according to claim 11, characterized in that The responding to the automatic marking instruction includes: responding to the automatic marking instruction when the automatic marking control is unlocked; wherein the unlocking condition of the automatic marking control is: the first training image has been marked and the second training image to be marked is displayed; The responding to the label appending instruction includes: responding to the label appending instruction when the smart appending control is unlocked; wherein the unlocking condition of the smart appending control is: there are a target box and a category label marked by the user on the first training image.

13. The method according to claim 1, wherein The method further comprises: When a modification instruction for any target frame is received, modifying the position, size, and / or corresponding category label of the target frame based on the modification instruction; When a deletion instruction is received for any target box, the target box and the corresponding category label are removed from the training image; When a new annotation added by the user for a training image is received, the target box and category label added by the user are displayed on the training image.

14. A method for labeling an image, characterized in that: Applied to electronic equipment, the method includes: Displaying an interactive interface for annotating training images; the interactive interface includes: an image display area, an option for adding training images, an option for annotating images, and a "Train" button; wherein the options for annotating images include: "Rectangular Selection" and "Smart Append"; After the user clicks the option to add a training image, a training image to be labeled selected by the user is displayed in the image display area; After the user clicks "rectangular box selection", the user's annotations for the training image are received, and the target boxes and category labels of different target objects annotated by the user are displayed on the training image; wherein each category label annotates at least one target object; In response to a label append instruction generated by the user clicking "Smart Append," displaying on the training image: a target frame and a category label of a target object in the training image; wherein the target frame and category label of the target object in the training image are obtained by inputting the training image into an object detection model trained using the training image annotated by the user; Based on the user's selection, the target box and category label of the target object to be annotated are corrected and adjusted; After the user clicks the "Train" button, the target detection model is trained using the labeled training images to obtain a trained target detection model.

15. The method according to claim 14, characterized in that There are multiple training images to be labeled; after executing the label appending instruction for the first training image, the method further includes: The next training image to be labeled is displayed in sequence, and the label appending instruction is executed for the training image until all training images are labeled.

16. The method according to claim 14, characterized in that After executing the label appending instruction for the first training image, the method further includes: Display the next training image to be labeled in sequence; In response to the automatic labeling instruction, displaying on the next training image: a target bounding box and a category label of a target object in the training image; wherein the target bounding box and the category label of the target object in the training image are obtained by inputting the training image into an object detection model trained using the labeled training image; Return to the step of sequentially displaying the next training image to be labeled until all training images are labeled.

17. A model training method, characterized in that: Applied to electronic equipment, the method includes: Obtaining a user-determined training image; Applying the image annotation method according to any one of claims 1 to 16 to annotate the training image; In response to the model training instruction, the target detection model is trained using the labeled training images to obtain a trained target detection model.

18. An image annotation device, characterized in that: Applied to electronic equipment, the device comprises: a display module configured to display an interactive interface for annotating training images; the interactive interface comprising an image display area, an option for adding a training image, an option for annotating an image, and a "train" button; after a user clicks the option to add a training image, a first training image to be annotated selected by the user is displayed in the image display area; a label receiving module, configured to receive the user's labeling for the first training image after the user clicks the labeling option, and display the target boxes and category labels of different target objects labeled by the user on the first training image; wherein each category label labels at least one target object; An automatic labeling response module is configured to display a second training image to be labeled selected by the user in the image display area; in response to an automatic labeling instruction generated by the user clicking an automatic labeling control, display on the second training image: target boxes and category labels of target objects in the second training image; based on the user's selection, display on the second training image the target boxes and category labels of all target objects that have been corrected and adjusted by the user; wherein the target boxes and category labels of the target objects in the second training image are obtained by inputting the second training image into a target detection model trained using the labeled first training image; the labeled first training image includes the target boxes and category labels labeled by the user; The model training instruction response module is used to train the target detection model using the labeled training images after the user clicks the "Train" button to obtain a trained target detection model.

19. An image annotation device, characterized in that: Applied to electronic equipment, the device comprises: A training image display module, configured to display an interactive interface for annotating training images; the interactive interface includes an image display area, an option for adding training images, an option for annotating images, and a "Train" button; the options for annotating images include "Rectangular Selection" and "Smart Append"; After the user clicks the option to add a training image, a training image to be labeled selected by the user is displayed in the image display area; A user annotation receiving module is configured to receive user annotations for the training image after the user clicks "rectangular box selection" and display target boxes and category labels of different target objects annotated by the user on the training image; wherein each category label annotates at least one target object; a label appending instruction response module, configured to respond to a label appending instruction generated by a user clicking "Smart Append" by displaying on the training image: a target frame and a category label of a target object in the training image; and, based on a user selection, modifying and adjusting the target frame and category label of the target object annotated this time; wherein the target frame and category label of the target object in the training image are obtained by inputting the training image into an object detection model trained using the user-annotated training image; The model training instruction response module is used to train the target detection model using the labeled training images after the user clicks the "Train" button to obtain a trained target detection model.

20. A model training device, characterized in that: Applied to electronic equipment, the device comprises: A training image acquisition module, used to obtain a training image determined by a user; a training image annotation module, configured to apply the image annotation method according to any one of claims 1 to 16 to annotate the training image; The model training instruction response module is used to respond to the model training instruction and use the labeled training images to train the target detection model to obtain a trained target detection model.

21. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the method of any one of claims 1 to 13, any one of claims 14 to 16, or claim 17 when executing a program stored in a memory.

22. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 13, any one of claims 14 to 16, or claim 17 is implemented.

23. A computer program product comprising instructions, characterized in that When the computer program product is run on a computer, the computer is caused to execute the method according to any one of claims 1 to 13, any one of claims 14 to 16, or claim 17.

Citation Information

Patent Citations

  • Picture labeling method and device, storage medium and electronic device

    CN110610169A

  • Data annotation method and device, electronic equipment and storage medium

    CN111680753A

  • Sample set construction method, question and answer model training method, question and answer processing method, request processing method and task platform

    CN119904880A