Image mask labeling method and device, storage medium and program product

By fine-tuning and merging the perspective image of the machine learning model, the problem of poor X-ray image processing of AI models is solved, and the efficiency and accuracy of pixel-level mask annotation are improved.

CN120298828APending Publication Date: 2025-07-11NUCTECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510434203.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

In the prior art, AI models have poor processing effects on X-ray images, resulting in inefficient mask labeling at the pixel level.

Method used

By fine-tuning the machine learning model with perspective images, the combined images are generated and the mask annotation information is written into the training data set, and the fine-tuning model is used for mask annotation.

Benefits of technology

Improves the efficiency and accuracy of mask annotation, reduces the calculation amount and reduces false positives.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298828A_ABST
    Figure CN120298828A_ABST
Patent Text Reader

Abstract

The invention provides an image mask labeling method and device, a storage medium and a program product. The image mask labeling method comprises the steps that mask labeling is carried out on a first perspective image to obtain mask labeling information, and the first perspective image comprises a predetermined suspect; the first perspective image and a second perspective image which is not subjected to mask labeling are combined to generate a combined image, and the second perspective image does not comprise any suspect; writing the area image of the area where the predetermined suspected object in the combined image is located and the mask labeling information into a training data set; performing fine tuning on the machine learning model by using the training data set; and performing mask labeling on each to-be-labeled perspective image in the to-be-labeled perspective image set by using the fine-tuned machine learning model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence, and particularly to an image mask annotation method and apparatus, a storage medium, and a program product. Background Art

[0002] Data annotation is a crucial link in the field of artificial intelligence. The annotation of rectangular boxes is relatively simple, but for pixel-level mask annotation, more time and labor costs are required to complete. In this case, the technology of collaborative annotation between humans and AI (Artificial Intelligence) models has emerged. This technology first uses a pre-trained AI model to pre-annotate an image, and for the inaccurately annotated and error-prone areas, the edges of the target object are precisely adjusted through a series of interaction information, so as to achieve refined annotation. Summary of the Invention

[0003] The inventors noticed that in the related technology of collaborative annotation between humans and AI models, the training images used for pre-training the AI model are visible light images, resulting in poor processing results of the AI model for X-ray images.

[0004] Accordingly, the present disclosure provides an image mask annotation method, which fine-tunes a machine learning model by using a perspective image, so that the fine-tuned machine learning model can effectively perform mask annotation on the perspective image, and effectively improve the mask annotation efficiency.

[0005] In a first aspect of the present disclosure, there is provided an image mask annotation method, including: performing mask annotation on a first perspective image to obtain mask annotation information, where the first perspective image includes a predetermined suspect; merging the first perspective image with a second perspective image that has not been mask-annotated to generate a merged image, where the second perspective image does not include any suspects; writing the region image of the region where the predetermined suspect is located in the merged image and the mask annotation information into a training data set; fine-tuning a machine learning model by using the training data set; and performing mask annotation on each to-be-annotated perspective image in a to-be-annotated perspective image set by using the fine-tuned machine learning model.

[0006] In some embodiments, the training data set and the to-be-annotated perspective image set are updated by moving the to-be-annotated perspective images with correct mask annotation results and the corresponding mask annotation results in the to-be-annotated perspective image set into the training data set; and the above steps of fine-tuning the machine learning model by using the training data set are repeatedly executed until the to-be-annotated perspective image set is an empty set.

[0007] In some embodiments, updating the training data set and the perspective images to be annotated includes: detecting whether the mask annotation result of each perspective image to be annotated in the perspective images to be annotated is correct; if the mask annotation result of the i-th perspective image to be annotated is correct, writing the i-th perspective image to be annotated and the corresponding mask annotation result into the training data set to update the training data set, , where N is the total number of perspective images to be annotated in the perspective images to be annotated; deleting the i-th perspective image to be annotated from the perspective images to be annotated to update the perspective images to be annotated.

[0008] In some embodiments, repeatedly performing fine-tuning of the machine learning model using the training data set includes: when each perspective image to be annotated in the perspective images to be annotated has been detected and the perspective images to be annotated are not empty, repeatedly performing fine-tuning of the machine learning model using the training data set until the perspective images to be annotated are an empty set.

[0009] In some embodiments, updating the training data set and the perspective images to be annotated includes: sorting all the perspective images to be annotated in the perspective images to be annotated in descending order of the confidence of the mask annotation result of each perspective image provided by the machine learning model; grouping all the perspective images to be annotated according to the sorting result to obtain multiple groups of perspective images to be annotated; sequentially detecting whether the mask annotation result of each perspective image to be annotated in each group of perspective images to be annotated is correct; if the mask annotation results of k perspective images to be annotated in the m-th group of perspective images to be annotated are correct, selecting the k perspective images to be annotated, , , where M is the total number of groups of perspective images to be annotated, and Km is the total number of perspective images to be annotated in the m-th group of perspective images to be annotated; in response to the user's confirmation operation, writing the k perspective images to be annotated and the corresponding mask annotation results into the training data set; deleting the k perspective images to be annotated from the perspective images to be annotated to update the perspective images to be annotated.

[0010] In some embodiments, repeatedly performing fine-tuning of the machine learning model using the training data set includes: when all the groups of perspective images to be annotated have been detected and the perspective images to be annotated are not empty, repeatedly performing fine-tuning of the machine learning model using the training data set until the perspective images to be annotated are an empty set.

[0011] In some embodiments, when the set of perspective images to be labeled is not empty, detect the number of perspective images to be labeled in the set; if the number of perspective images to be labeled is greater than a predetermined threshold, randomly select multiple perspective images to be labeled from the set of perspective images to be labeled as multiple images to be processed; manually correct the mask annotation results of each image to be processed among the multiple images to be processed; write each image to be processed and the corresponding corrected mask annotation result into the training dataset; delete the multiple images to be processed from the set of perspective images to be labeled to update the set of perspective images to be labeled.

[0012] In some embodiments, if the number of perspective images to be labeled is not greater than the predetermined threshold, manually correct the mask annotation results of each perspective image to be labeled in the set of perspective images to be labeled.

[0013] In some embodiments, use the target region image in each original perspective image in the original perspective image set as the perspective image to be labeled to generate the set of perspective images to be labeled.

[0014] In some embodiments, the target region image in each original perspective image is obtained through a ground truth box or a candidate box.

[0015] In a second aspect of the present disclosure, there is provided an image mask annotation device, including: a preprocessing module configured to perform mask annotation on a first perspective image to obtain mask annotation information, where the first perspective image includes a predetermined suspect; a merging module configured to merge the first perspective image with a second perspective image that has not been mask-annotated to generate a merged image, where the second perspective image does not include any suspects; a training module configured to write the region image of the area where the predetermined suspect is located in the merged image and the mask annotation information into a training dataset, and fine-tune a machine learning model using the training dataset; and an annotation management module configured to perform mask annotation on each perspective image to be labeled in the set of perspective images to be labeled using the fine-tuned machine learning model.

[0016] In a third aspect of the present disclosure, there is provided an image mask annotation device, including: a memory; a processor coupled to the memory, the processor being configured to execute based on instructions stored in the memory to implement the image mask annotation method as described in any one of the above embodiments.

[0017] According to a fourth aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium, where the computer-readable storage medium stores computer instructions, and when the instructions are executed by a processor, the image mask annotation method as described in any one of the above embodiments is implemented.

[0018] According to a fifth aspect of the embodiments of the present disclosure, there is provided a computer program product including computer instructions, where when the computer instructions are executed by a processor, the image mask annotation method described in any of the above embodiments is implemented.

[0019] Other features and advantages of the present disclosure will become clear from the following detailed description of the exemplary embodiments of the present disclosure with reference to the accompanying drawings. Description of the Drawings

[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0021] Figure 1 Schematic flowchart of the image mask annotation method according to an embodiment of the present disclosure;

[0022] Figure 2 Schematic diagram of a TIP image according to an embodiment of the present disclosure;

[0023] Figure 3 Schematic diagram of a perspective image without a suspect according to an embodiment of the present disclosure;

[0024] Figure 4 Schematic diagram of a merged image according to an embodiment of the present disclosure;

[0025] Figure 5 Schematic diagram of an original perspective image according to an embodiment of the present disclosure;

[0026] Figure 6 Schematic flowchart of the image mask annotation method according to another embodiment of the present disclosure;

[0027] Figure 7 Schematic structural diagram of an image mask annotation device according to an embodiment of the present disclosure;

[0028] Figure 8 Schematic structural diagram of an image mask annotation device according to another embodiment of the present disclosure. Detailed Embodiments

[0029] Next, the technical solutions in the embodiments of the present disclosure will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all the embodiments. The following description of at least one exemplary embodiment is actually illustrative only and in no way limits the present disclosure, its application or use. Based on the embodiments in the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present disclosure.

[0030] Unless otherwise specifically stated, the relative arrangements, numerical expressions and values of the components and steps set forth in these embodiments do not limit the scope of the present disclosure.

[0031] At the same time, it should be understood that, for the sake of convenience of description, the sizes of the various parts shown in the accompanying drawings are not drawn according to the actual proportional relationship.

[0032] For technologies, methods and devices known to those of ordinary skill in the relevant art, they may not be discussed in detail, but under appropriate circumstances, the technologies, methods and devices should be regarded as part of the authorization specification.

[0033] In all the examples shown and discussed here, any specific value should be interpreted as merely exemplary, rather than as a limitation. Therefore, other examples of the exemplary embodiments may have different values.

[0034] It should be noted that: like reference numerals and letters denote like items in the following accompanying drawings. Therefore, once an item is defined in one of the accompanying drawings, it does not need to be further discussed in the subsequent accompanying drawings.

[0035] Figure 1 It is a schematic flow chart of an image mask annotation method according to an embodiment of the present disclosure. In some embodiments, the following image mask annotation method is executed by an image mask annotation device and includes steps 11-15.

[0036] In step 11, mask annotation is performed on a first perspective image to obtain mask annotation information, where the first perspective image includes a predetermined suspect.

[0037] It should be noted here that the complexity of the first perspective image is much lower than that of the perspective images collected under normal circumstances. For example, the first perspective image only includes a single suspect and the image background is white, so that traditional image processing algorithms can be used to accurately perform mask annotation on the first perspective image.

[0038] For example, the first perspective image can also be referred to as a TIP (Threat Image Projection) image.

[0039] For example, in Figure 2 In a TIP image shown, only one suspect is included, and the background of the TIP is white, thus facilitating accurate mask annotation for the TIP image.

[0040] It should be noted that mask annotation refers to generating a pixel-level mask for the target object, which can more precisely represent the shape and boundary of the target object.

[0041] In step 12, the first perspective image is merged with the second perspective image without mask annotation to generate a merged image, where no suspect is included in the second perspective image.

[0042] For example, the second perspective image is as Figure 3 shown. After merging Figure 2 and Figure 3 , the obtained merged image is as Figure 4 shown.

[0043] It should be noted here that since the second perspective image has no mask annotation and does not include any suspect, after merging Figure 2 and Figure 3 , it will not affect the subsequent processing of the merged image.

[0044] In step 13, the region image and mask annotation information of the region where the predetermined suspect is located in the merged image are written into the training dataset.

[0045] In step 14, the machine learning model is fine-tuned using the training dataset.

[0046] It should be noted here that fine-tuning refers to performing targeted training on the machine learning model, which can significantly improve the performance of the machine learning model on specific tasks.

[0047] In step 15, each perspective image to be annotated in the set of perspective images to be annotated is mask-annotated using the fine-tuned machine learning model.

[0048] In some embodiments, the target region image in each original perspective image in the original set of perspective images is used as the perspective image to be annotated to generate a set of perspective images to be annotated.

[0049] For example, the target region image in each original perspective image is obtained through a Ground Truth Box or a proposal.

[0050] For example, in Figure 5 the shown original perspective image, a folding knife is hidden in the handbag.

[0051] It should be noted here that if mask annotation is performed on the entire original perspective image, the computational complexity of mask annotation will increase, and at the same time, the false alarm of full-image annotation information will also increase. To solve this problem, by using the target region image in the original perspective image as the perspective image to be annotated, the computational complexity of mask annotation can be effectively reduced, the occurrence of false alarms of full-image annotation information can be effectively avoided, and at the same time, the review efficiency of annotation information can also be effectively improved.

[0052] In the image mask annotation method provided in the above embodiments of the present disclosure, by using the perspective image to fine-tune the machine learning model, the fine-tuned machine learning model can effectively perform mask annotation on the perspective image, effectively improving the mask annotation efficiency.

[0053] In some embodiments, by moving the perspective image to be annotated with a correct mask annotation result and the corresponding mask annotation result in the set of perspective images to be annotated into the training data set, the training data set and the set of perspective images to be annotated are updated. Next, the step of fine-tuning the machine learning model using the training data set is repeatedly executed until the set of perspective images to be annotated becomes an empty set.

[0054] In some embodiments, the steps of updating the training data set and the set of perspective images to be annotated include the following content.

[0055] 1) Detect whether the mask annotation result of each perspective image to be annotated in the set of perspective images to be annotated is correct.

[0056] For example, a conventional annotation verification tool can be used to detect whether the mask annotation result of each perspective image to be annotated is correct.

[0057] 2) If the mask annotation result of the i-th perspective image to be annotated is correct, write the i-th perspective image to be annotated and the corresponding mask annotation result into the training data set to update the training data set, , where N is the total number of perspective images to be annotated in the set of perspective images to be annotated.

[0058] 3) Delete the i-th perspective image to be annotated from the set of perspective images to be annotated to update the set of perspective images to be annotated.

[0059] That is to say, the training data set and the set of perspective images to be annotated are updated according to the mask annotation result of each perspective image to be annotated.

[0060] In this case, when each perspective image to be annotated in the set of perspective images to be annotated is detected and the set of perspective images to be annotated is not an empty set, the step of fine-tuning the machine learning model using the training data set is repeatedly executed until the set of perspective images to be annotated becomes an empty set.

[0061] For example, if there are 1000 perspective images to be annotated in the perspective image set to be annotated, and if the mask annotation results of 400 perspective images to be annotated are correct, then these 400 perspective images to be annotated and the corresponding mask annotation results are sequentially written into the training data set to update the training data set. At the same time, these 400 perspective images to be annotated are deleted from the perspective image set to be annotated to update the perspective image set to be annotated. Next, the step of fine-tuning the machine learning model using the training data set is repeatedly executed until the perspective image set to be annotated becomes an empty set.

[0062] In some embodiments, the steps of updating the training data set and the perspective image set to be annotated include the following.

[0063] 1) Sort all the perspective images to be annotated in the perspective image set to be annotated in descending order of the confidence of the mask annotation results provided by the machine learning model for each perspective image to be annotated.

[0064] 2) Group all the perspective images to be annotated according to the sorting result to obtain multiple groups of perspective images to be annotated.

[0065] 3) Sequentially detect whether the mask annotation results of each perspective image to be annotated in each group of perspective images to be annotated are correct.

[0066] 4) If the mask annotation results of k perspective images to be annotated in the m-th group of perspective images to be annotated are correct, then select the k perspective images to be annotated, , , where M is the total number of groups of perspective images to be annotated, and Km is the total number of perspective images to be annotated in the m-th group of perspective images to be annotated.

[0067] 5) In response to the user's confirmation operation, write the k perspective images to be annotated and the corresponding mask annotation results into the training data set.

[0068] 6) Delete the k perspective images to be annotated from the perspective image set to be annotated to update the perspective image set to be annotated.

[0069] In this case, when multiple groups of perspective images to be annotated have been detected and the perspective image set to be annotated is not an empty set, the step of fine-tuning the machine learning model using the training data set is repeatedly executed until the perspective image set to be annotated becomes an empty set.

[0070] For example, there are 1000 perspective images to be annotated in the perspective image set to be annotated. Sort these 1000 perspective images in descending order of the confidence of the mask annotation results, and group these 1000 perspective images. Each group of perspective images to be annotated includes 100 perspective images to be annotated.

[0071] For the first group of perspective images to be labeled, the confidence levels of the mask annotation results of the 100 perspective images to be labeled included in this group are all very high. Through detection, the mask annotation results of the 100 perspective images to be labeled included in this group are all correct. In this case, by selecting these 100 perspective images to be labeled and according to the user's confirmation operation, write these 100 perspective images to be labeled and the corresponding mask annotation results into the training dataset to update the training dataset, and delete these 100 perspective images to be labeled from the perspective image set to be labeled to update the perspective image set to be labeled.

[0072] For the second group of perspective images to be labeled, the confidence levels of the mask annotation results of the 100 perspective images to be labeled included in this group are also very high. Through detection, the mask annotation results of 95 perspective images to be labeled included in this group are all correct. In this case, by selecting these 95 perspective images to be labeled and according to the user's confirmation operation, write these 95 perspective images to be labeled and the corresponding mask annotation results into the training dataset to update the training dataset, and delete these 95 perspective images to be labeled from the perspective image set to be labeled to update the perspective image set to be labeled.

[0073] Thus, through one-key operation, the processing of multiple perspective images to be labeled can be achieved, thereby effectively improving the processing efficiency.

[0074] In some embodiments, when the perspective image set to be labeled is not empty, the following operations are further performed.

[0075] 1) Detect the number of images to be labeled in the perspective image set to be labeled.

[0076] 2) If the number of perspective images to be labeled is greater than a predetermined threshold, randomly select multiple perspective images to be labeled in the perspective image set to be labeled as multiple images to be processed.

[0077] 3) Manually correct the mask annotation results of each image to be processed among the multiple images to be processed.

[0078] 4) Write each image to be processed and the corresponding corrected mask annotation results into the training dataset.

[0079] 5) Delete the multiple images to be processed from the perspective image set to be labeled to update the perspective image set to be labeled.

[0080] That is to say, when there are many perspective images to be labeled in the perspective image set to be labeled, randomly select some perspective images to be labeled, manually correct the mask annotation results of the selected images to be processed, and then use the manually corrected mask annotation results to participate in the further fine-tuning of the machine learning model, thereby further improving the mask annotation accuracy of the machine learning model.

[0081] In some embodiments, if the number of perspective images to be labeled is not greater than a predetermined threshold, the mask annotation results of each perspective image to be labeled in the set of perspective images to be labeled are manually corrected.

[0082] For example, if after multiple rounds of loops, the number of perspective images to be labeled in the set of perspective images to be labeled is small, for example, not exceeding 5, in this case, the mask annotation results of these perspective images to be labeled can be directly manually corrected, which helps to quickly complete the image annotation work.

[0083] Figure 6 It is a schematic flowchart of an image mask annotation method according to another embodiment of the present disclosure. In some embodiments, the following image mask annotation method is executed by an image mask annotation device, including steps 61-68.

[0084] In step 61, mask annotation is performed on a first perspective image to obtain mask annotation information, where the first perspective image includes a predetermined suspect.

[0085] For example, the first perspective image is a TIP image.

[0086] In step 62, a second perspective image that has not been mask-annotated is obtained, where the second perspective image does not include any suspects.

[0087] In step 63, the first perspective image and the second perspective image are merged to generate a merged image.

[0088] In step 64, the region image of the region where the predetermined suspect is located in the merged image and the mask annotation information are written into the training data set.

[0089] In step 65, the machine learning model is fine-tuned using the training data set.

[0090] In step 66, using the fine-tuned machine learning model, mask annotation is performed on each perspective image to be labeled in the set of perspective images to be labeled.

[0091] In step 67, it is detected whether the mask annotation result of each perspective image to be labeled in the set of perspective images to be labeled is correct.

[0092] According to the detection result, return to step 64, that is, the perspective image to be labeled with a correct mask annotation result in the set of perspective images to be labeled and the corresponding mask annotation result are moved into the training data set, and the training data set and the set of perspective images to be labeled are updated.

[0093] It should be noted that the above-mentioned embodiment-involved scheme of detecting the mask annotation results of each perspective image to be annotated one by one can be adopted, or the above-mentioned embodiment-involved scheme of grouping multiple perspective images to be annotated included in the perspective image set to be annotated and detecting each group accordingly can be adopted.

[0094] In step 68, when the perspective image set to be annotated is not empty, randomly select some perspective images to be annotated, manually correct the mask annotation results of the selected images to be processed, and then return to step 64, that is, write each image to be processed and the corresponding corrected mask annotation results into the training dataset, and delete multiple images to be processed from the perspective image set to be annotated to update the perspective image set to be annotated.

[0095] Through the above processing, the fine-tuned machine learning model can effectively perform mask annotation on perspective images, effectively improving the mask annotation efficiency.

[0096] Figure 7 It is a schematic structural diagram of an image mask annotation device according to an embodiment of the present disclosure.

[0097] As Figure 7 shown, the image mask annotation device includes a preprocessing module 71, a merging module 72, a training module 73, and an annotation management module 74.

[0098] The preprocessing module 71 is configured to perform mask annotation on a first perspective image to obtain mask annotation information, where the first perspective image includes a predetermined suspect.

[0099] It should be noted that the complexity of the first perspective image is much lower than that of the perspective images collected under normal circumstances. For example, the first perspective image only includes a single suspect, and the image background is white, so that traditional image processing algorithms can be used to accurately perform mask annotation on the first perspective image.

[0100] For example, the first perspective image can also be referred to as a TIP image.

[0101] The merging module 72 is configured to merge the first perspective image with a second perspective image that has not been mask-annotated to generate a merged image, where the second perspective image does not include any suspects.

[0102] It should be noted here that since the second perspective image has not been mask-annotated and does not include any suspects, after merging the first perspective image and the second perspective image, it will not affect the subsequent processing of the merged image.

[0103] The training module 73 is configured to write the region image and mask annotation information of the region where the predetermined suspect is located in the merged image into the training dataset, and use the training dataset to fine-tune the machine learning model.

[0104] The annotation management module 74 is configured to perform mask annotation on each perspective image to be annotated in the perspective image set to be annotated by using the fine-tuned machine learning model.

[0105] In some embodiments, the target region image in each original perspective image in the original perspective image set is used as the perspective image to be annotated, so as to generate the perspective image set to be annotated.

[0106] For example, the target region image in each original perspective image is obtained through a ground truth box or a candidate box.

[0107] In some embodiments, the annotation management module 74 updates the training dataset and the perspective image set to be annotated by moving the perspective image to be annotated with the correct mask annotation result in the perspective image set to be annotated and the corresponding mask annotation result into the training dataset. Next, the training module 73 repeats the operation of fine-tuning the machine learning model by using the training dataset until the perspective image set to be annotated is an empty set.

[0108] In some embodiments, the operation of the annotation management module 74 to update the training dataset and the perspective image set to be annotated includes the following contents.

[0109] 1) Detect whether the mask annotation result of each perspective image to be annotated in the perspective image set to be annotated is correct.

[0110] For example, a conventional annotation verification tool can be used to detect whether the mask annotation result of each perspective image to be annotated is correct.

[0111] 2) If the mask annotation result of the i-th perspective image to be annotated is correct, write the i-th perspective image to be annotated and the corresponding mask annotation result into the training dataset to update the training dataset, , where N is the total number of perspective images to be annotated in the perspective image set to be annotated.

[0112] 3) Delete the i-th perspective image to be annotated from the perspective image set to be annotated to update the perspective image set to be annotated.

[0113] That is, the training dataset and the perspective image set to be annotated are updated according to the mask annotation result of each perspective image to be annotated.

[0114] In this case, when the annotation management module 74 has detected each to-be-annotated perspective image in the to-be-annotated perspective image set and the to-be-annotated perspective image set is not empty, the training module 73 repeatedly performs fine-tuning of the machine learning model using the training data set until the to-be-annotated perspective image set becomes an empty set.

[0115] In some embodiments, the operations of the annotation management module 74 for updating the training data set and the to-be-annotated perspective image set include the following.

[0116] 1) Sort all the to-be-annotated perspective images in the to-be-annotated perspective image set in descending order of the confidence level of the mask annotation result provided by the machine learning model for each to-be-annotated perspective image.

[0117] 2) Group all the to-be-annotated perspective images according to the sorting result to obtain multiple groups of to-be-annotated perspective images.

[0118] 3) Detect in sequence whether the mask annotation result of each to-be-annotated perspective image in each group of to-be-annotated perspective images is correct.

[0119] 4) If the mask annotation results of the k to-be-annotated perspective images in the m-th group of to-be-annotated perspective images are correct, select the k to-be-annotated perspective images, , , where M is the total number of groups of to-be-annotated perspective images, and Km is the total number of to-be-annotated perspective images in the m-th group of to-be-annotated perspective images.

[0120] 5) In response to the user's confirmation operation, write the k to-be-annotated perspective images and the corresponding mask annotation results into the training data set.

[0121] 6) Delete the k to-be-annotated perspective images from the to-be-annotated perspective image set to update the to-be-annotated perspective image set.

[0122] In this case, when the annotation management module 74 has detected multiple groups of to-be-annotated perspective images and the to-be-annotated perspective image set is not empty, the training module 73 repeatedly performs fine-tuning of the machine learning model using the training data set until the to-be-annotated perspective image set becomes an empty set.

[0123] In some embodiments, when the to-be-annotated perspective image set is not empty, the annotation management module 74 further performs the following operations.

[0124] 1) Detect the number of to-be-annotated images in the to-be-annotated perspective image set.

[0125] 2) If the number of to-be-annotated perspective images is greater than a predetermined threshold, randomly select multiple to-be-annotated perspective images from the to-be-annotated perspective image set as multiple to-be-processed images.

[0126] 3) Manually correct the mask annotation results of each image to be processed among multiple images to be processed.

[0127] 4) Write each image to be processed and the corresponding corrected mask annotation results into the training dataset.

[0128] 5) Delete multiple images to be processed from the set of perspective images to be annotated to update the set of perspective images to be annotated.

[0129] That is to say, when there are many perspective images to be annotated in the set of perspective images to be annotated, the annotation management module 74 randomly selects some perspective images to be annotated, and manually corrects the mask annotation results of the selected images to be processed. Then, the training module 73 uses the manually corrected mask annotation results to participate in the further fine-tuning of the machine learning model, thereby further improving the mask annotation accuracy of the machine learning model.

[0130] In some embodiments, if the number of perspective images to be annotated is not greater than a predetermined threshold, the annotation management module 74 manually corrects the mask annotation results of each perspective image to be annotated in the set of perspective images to be annotated.

[0131] Figure 8 It is a schematic structural diagram of an image mask annotation device according to another embodiment of the present disclosure;

[0132] As Figure 8 shown, the image mask annotation device 80 is presented in the form of a general-purpose computing device. The image mask annotation device 80 includes a memory 81, a processor 82, and a bus 83 connecting different system components.

[0133] The memory 81 may include, for example, a system memory and a non-volatile storage medium. The system memory stores, for example, an operating system, application programs, a boot loader, and other programs. The system memory may include a volatile storage medium, such as a random access memory (RAM) and / or a cache memory. The non-volatile storage medium stores, for example, instructions corresponding to at least one embodiment of the image mask annotation method in execution. The non-volatile storage medium includes, but is not limited to, a disk memory, an optical memory, a flash memory, etc.

[0134] The processor 82 may be implemented in the form of a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete hardware components such as discrete gates or transistors. Correspondingly, each module, such as an acquisition module, a calculation module, and an adjustment module, may be implemented by a central processing unit (CPU) running instructions in the memory to execute corresponding steps, or may be implemented by a dedicated circuit executing the corresponding steps.

[0135] For example, the processor 82 is configured to execute, based on instructions stored in the memory, a method involved in any one of the embodiments such as Figure 1 or Figure 6 .

[0136] The bus 83 can use any bus structure among a variety of bus structures. For example, the bus structure includes but is not limited to an Industry Standard Architecture (ISA) bus, a Micro Channel Architecture (MCA) bus, and a Peripheral Component Interconnect (PCI) bus.

[0137] These interfaces 84, 85, 86 of the image mask annotation device 80 and between the memory 81 and the processor 82 can be connected through the bus 83. The input / output interface 84 can provide a connection interface for input / output devices such as a display, a mouse, and a keyboard. The network interface 85 provides a connection interface for various networking devices. The storage interface 86 provides a connection interface for external storage devices such as a floppy disk, a USB flash drive, and an SD card.

[0138] Here, various aspects of the present disclosure have been described with reference to the flowcharts and / or block diagrams of methods, apparatuses, and computer program products according to embodiments of the present disclosure. It should be understood that each box of the flowcharts and / or block diagrams, and combinations of the boxes, can be implemented by computer-readable program instructions.

[0139] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable devices to generate a machine, such that the device implementing the functions specified in one or more boxes in the flowcharts and / or block diagrams is generated by executing the instructions by the processor.

[0140] These computer-readable program instructions can also be stored in a computer-readable memory, and these instructions cause the computer to work in a specific manner, thereby generating a manufactured article including instructions for implementing the functions specified in one or more boxes in the flowcharts and / or block diagrams.

[0141] The present disclosure can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects.

[0142] The present disclosure also provides a computer-readable storage medium, in which computer instructions are stored, and when the instructions are executed by a processor, a method involved in any one of the embodiments such as Figure 1 or Figure 6 is implemented.

[0143] The present disclosure also provides a computer program product, including computer instructions, and when the computer instructions are executed by a processor, a method involved in any one of the embodiments such as Figure 1 or Figure 6 is implemented.

[0144] By implementing the above embodiments of the present disclosure, it is possible to enable the fine-tuned machine learning model to effectively perform mask annotation on perspective images, effectively improving the mask annotation efficiency. For example, by implementing the above technical solutions of the present disclosure, the accuracy of the pre-labeled mask can reach 97.6%, and the recall rate of the ground-truth mask can reach 86.3%.

[0145] In some embodiments, the functional units described above may be implemented as a general-purpose processor, a programmable logic controller (PLC), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, or any suitable combination thereof for performing the functions described in the present disclosure.

[0146] Those of ordinary skill in the art can understand that all or part of the steps for implementing the above embodiments can be completed by hardware or by a program instructing the relevant hardware. The program can be stored in a computer-readable storage medium. The storage medium mentioned above can be a read-only memory, a magnetic disk, an optical disk, or the like.

[0147] The description of the present disclosure has been presented for purposes of illustration and description, and is not intended to be exhaustive or to limit the present disclosure to the form disclosed. Many modifications and variations will be apparent to those of ordinary skill in the art. The embodiments were chosen and described in order to best explain the principles of the present disclosure and its practical application, and to enable those of ordinary skill in the art to understand the present disclosure and design various embodiments with various modifications suitable for specific purposes.

Claims

1. An image mask annotation method, comprising: Performing mask annotation on a first perspective image to obtain mask annotation information, wherein the first perspective image includes a predetermined suspect; Merging the first perspective image with a second perspective image without mask annotation to generate a merged image, wherein the second perspective image does not include any suspects; Writing the region image of the region where the predetermined suspect is located in the merged image and the mask annotation information into a training data set; Fine-tuning a machine learning model using the training data set; Performing mask annotation on each perspective image to be annotated in a set of perspective images to be annotated using the fine-tuned machine learning model.

2. The image mask annotation method according to claim 1, further comprising: Updating the training data set and the set of perspective images to be annotated by moving the perspective images to be annotated with correct mask annotation results in the set of perspective images to be annotated and the corresponding mask annotation results into the training data set; Repeatedly performing fine-tuning of the machine learning model using the training data set until the set of perspective images to be annotated is an empty set.

3. The image mask annotation method according to claim 2, wherein The updating the training data set and the set of perspective images to be annotated includes: Detecting whether the mask annotation result of each perspective image to be annotated in the set of perspective images to be annotated is correct; If the mask annotation result of the i-th perspective image to be annotated is correct, write the i-th perspective image to be annotated and the corresponding mask annotation result into the training dataset to update the training dataset. , where N is the total number of perspective images to be annotated in the set of perspective images to be annotated; Deleting the i-th perspective image to be annotated from the set of perspective images to be annotated to update the set of perspective images to be annotated.

4. The image mask annotation method according to claim 3, wherein, The repeatedly performing fine-tuning of the machine learning model using the training data set includes: When each perspective image to be annotated in the set of perspective images to be annotated has been detected and the set of perspective images to be annotated is not an empty set, repeatedly performing fine-tuning of the machine learning model using the training data set until the set of perspective images to be annotated is an empty set.

5. The image mask annotation method according to claim 2, wherein, The updating the training data set and the set of perspective images to be annotated includes: Sorting all the perspective images to be annotated in the set of perspective images to be annotated in descending order of the confidence of the mask annotation results provided by the machine learning model for each perspective image to be annotated; Grouping all the perspective images to be annotated according to the sorting result to obtain multiple groups of perspective images to be annotated; Successively detecting whether the mask annotation result of each perspective image to be annotated in each group of perspective images to be annotated is correct; If the mask annotation results of k to-be-annotated perspective images in the m-th to-be-annotated perspective image group are correct, then select the k to-be-annotated perspective images. , , where M is the total number of to-be-annotated perspective image groups, and Km is the total number of to-be-annotated perspective images in the m-th to-be-annotated perspective image group; In response to a confirmation operation by the user, writing the k perspective images to be annotated and the corresponding mask annotation results into the training data set; Deleting the k perspective images to be annotated from the set of perspective images to be annotated to update the set of perspective images to be annotated.

6. The image mask annotation method according to claim 5, wherein, The repeatedly performing fine-tuning of the machine learning model using the training data set includes: When each of the multiple groups of perspective images to be annotated has been detected and the set of perspective images to be annotated is not an empty set, repeatedly performing fine-tuning of the machine learning model using the training data set until the set of perspective images to be annotated is an empty set.

7. The image mask annotation method according to claim 2, further comprising: When the set of perspective images to be labeled is not empty, detect the number of perspective images to be labeled in the set of perspective images to be labeled; If the number of perspective images to be labeled is greater than a predetermined threshold, randomly select a plurality of perspective images to be labeled from the set of perspective images to be labeled as a plurality of images to be processed; Manually correct the mask annotation results of each image to be processed among the plurality of images to be processed; Write each image to be processed and the corresponding corrected mask annotation result into the training data set; Delete the plurality of images to be processed from the set of perspective images to be labeled to update the set of perspective images to be labeled.

8. The image mask annotation method according to claim 7, further comprising: If the number of perspective images to be labeled is not greater than the predetermined threshold, manually correct the mask annotation results of each perspective image to be labeled in the set of perspective images to be labeled.

9. The image mask annotation method according to any one of claims 1-8, further comprising: Use the target region image in each original perspective image in the original perspective image set as the perspective image to be labeled to generate the set of perspective images to be labeled.

10. The image mask annotation method according to claim 9, wherein The target region image in each original perspective image is obtained through a ground truth box or a candidate box.

11. An image mask annotation device, comprising: A preprocessing module configured to perform mask annotation on a first perspective image to obtain mask annotation information, wherein the first perspective image includes a predetermined suspect; A merging module configured to merge the first perspective image with a second perspective image that has not been mask-annotated to generate a merged image, wherein the second perspective image does not include any suspects; A training module configured to write the region image of the area where the predetermined suspect is located in the merged image and the mask annotation information into a training data set, and fine-tune a machine learning model using the training data set; A labeling management module configured to perform mask annotation on each perspective image to be labeled in the set of perspective images to be labeled using the fine-tuned machine learning model.

12. An image mask annotation device, comprising: A memory; A processor coupled to the memory, the processor being configured to execute based on instructions stored in the memory to implement the image mask annotation method according to any one of claims 1-10.

13. A computer-readable storage medium, wherein, A computer-readable storage medium stores computer instructions, and when the instructions are executed by a processor, the image mask annotation method according to any one of claims 1-10 is implemented.

14. A computer program product, comprising computer instructions, wherein when the computer instructions are executed by a processor, the image mask annotation method according to any one of claims 1-10 is implemented.