Automated image annotation method and system

The automated image annotation method addresses the challenge of labeling uncommon images by using a neural network to compute feature vectors from user-selected objects in a training image, enabling accurate inference and annotation of similar objects in other images, thereby improving efficiency and accuracy in image annotation.

WO2025119487A1PCT designated stage expired Publication Date: 2025-06-12ROBERT BOSCH GMBH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2023/084818
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-08
Publication Date
2025-06-12

AI Technical Summary

Technical Problem

Existing automated image annotation methods struggle to label uncommon images, such as industrial or medical images, as they rely on pre-trained models, object geometry, or external data sources, which are inadequate for recognizing general objects without characteristic geometries or sufficient labeled examples.

Method used

The method involves providing a training image for user selection of objects, segmenting the image into patches, computing feature vectors using a neural network, and inferring object occurrences in sample images by comparing patch feature vectors. This approach allows for one-shot learning and multi-class labeling without requiring prior knowledge about the objects.

Benefits of technology

The method achieves high precision (83.8%) in annotating images by accurately identifying screws in industrial images, significantly reducing the time and effort required for manual annotation and demonstrating the potential for scalable and accurate automated image annotation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2023084818_12062025_PF_FP_ABST
    Figure EP2023084818_12062025_PF_FP_ABST
Patent Text Reader

Abstract

A method and system for automated image annotation. The method comprises the steps of providing a training image for user selection of one or more occurrences of at least one object within the image; segmenting the training image into a first plurality of patches; identifying one or more patches out of the first plurality of patches that correspond to respective ones of the user selected one or more occurrences of the at least one object; computing, using a neural network, a feature vector for each of the one or more patches out of the first plurality of patches; inferring occurrences of the at least one object in one or more sample images; and annotating each of the one or more sample images based on each inferred occurrence of the at least one object for display.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] AUTOMATED IMAGE ANNOTATION METHOD AND SYSTEM

[0002] FIELD OF INVENTION

[0003] The present invention relates broadly to an automated image annotation method and system, in particular to automated image annotation based on an image segmentation foundation model.

[0004] BACKGROUND

[0005] Any mention and / or discussion of prior art throughout the specification should not be considered, in any way, as an admission that this prior art is well known or forms part of common general knowledge in the field.

[0006] High-quality data annotation is the key to the successful application of deep learning models. Following the public release of ImageNet in 2009 - a large scale image recognition dataset with human labels - the best performing architectures increased from 71.8% in 2010 [1] to 95.5% top-5 accuracy in 2015 [2], far exceeding human performance. However, manual labelling is a tedious process, requiring significant amounts of both time and money [3], Not all organizations can afford such high overheads, which is why new methods are needed to alleviate the prohibitive cost of annotation.

[0007] Automated labelling usually comes down to image recognition. That is, one needs to understand an image in order to provide a label for it. But for this task, there is the added constraint of having no or few image-label pairs available. Most existing methods usually fall under one of the three following categories:

[0008] (1) Leveraging pre-trained models: By taking advantage of large neural networks pretrained on a variety of object categories, one can detect those same object categories in new sets of images. The models are most often pretrained in a supervised way and used in a zero-shot manner [5] (if the dataset used for pretraining is close enough to the downstream distribution).

[0009] (2) Using the object geometry: When the objects of interest have a peculiar geometry, one may try to utilize this in order to automatically generate labels. In the autonomous driving field, researchers have shown that poles are easily recognized in LiDAR range images due to their elongated shape [6], enabling to build potentially large sets of pseudo-labelled data. In automated disassembly, [7] uses Hough transforms to detect screw thanks to their circular shape. In both cases, the authors showed empirically that these pseudo-labels were strong enough to be used as training data for a neural network.

[0010] (3) Using external information sources: In some cases, one may have access to other types of information in order to label images. In [8], the authors use publicly available maps that contain the locations of poles and other types of objects in order to label images taken at these locations. The existing methods discussed above enable to label objects in a variety of images, but they fail to recognize items that are either uncommon or that do not have a characteristic geometry. For example, pre-trained neural networks are utterly unable to recognize anything outside of the training distribution. Since these models are usually trained on everyday objects (from datasets like ImageNet or COCO), these approaches are not well suited for industrial or medical images, which have a very different distribution. Although it is always possible to fine-tune the models to adapt them to the target distribution, this process requires labelled examples, which defies the purpose of automated labelling itself.

[0011] On the other hand, methods based on object geometry may be used for specific object shapes, but they fail to recognize more general objects. They also require some prior knowledge about the data in order to build a pipeline suited for each shape one wants to detect, which can lead to a lot of time spent on feature engineering.

[0012] Finally, depending on the application domain, there simply may not be any external data sources that one can exploit to facilitate annotation, making the methods described in item (3) above irrelevant in many applications.

[0013] All in all, this means that there is no existing method readily available that one could leverage in order to automatically label uncommon images (for example, industrial images).

[0014] Embodiments of the present invention seek to address at least one of the above problems.

[0015] SUMMARY

[0016] In accordance with a first aspect of the present invention, there is provided a method for automated image annotation comprising the steps of: providing a training image for user selection of one or more occurrences of at least one object within the image; segmenting the training image into a first plurality of patches; identifying one or more patches out of the first plurality of patches that correspond to respective ones of the user selected one or more occurrences of the at least one object; compute, using a neural network, a feature vector for each of the one or more patches out of the first plurality of patches; inferring occurrences of the at least one object in one or more sample images; and annotating each of the one or more sample images based on each inferred occurrence of the at least one object for display; wherein inferring the occurrences of the at least one object comprises: segmenting each sample image into a second plurality of patches; compute, using the neural network, a feature vector for each of the patches of the second plurality of patches; and inferring the occurrences of the at least one object based on a comparison of the feature vectors for each of the patches of the second plurality of patches with the feature vectors for each of the one or more patches out of the first plurality of patches.

[0017] In accordance with a second aspect of the present invention, there is provided a system for automated image annotation comprising: a display device for providing a training image for user selection of one or more occurrences of at least one object within the image and for displaying annotated versions of one or more sample images; a processing unit configured to: segmenting the training image into a first plurality of patches; identifying one or more patches out of the first plurality of patches that correspond to respective ones of the user selected one or more occurrences of the at least one object; compute, using a neural network, a feature vector for each of the one or more patches out of the first plurality of patches; inferring occurrences of the at least one object in the one or more sample images; and annotating each of the one or more sample images based on each inferred occurrence of the at least one object for display of the annotated versions of the one or more sample images; wherein inferring the occurrences of the at least one object comprises: segmenting each sample image into a second plurality of patches; compute, using the neural network, a feature vector for each of the patches of the second plurality of patches; and inferring the occurrences of the at least one object based on a comparison of the feature vectors for each of the patches of the second plurality of patches with the feature vectors for each of the one or more patches out of the first plurality of patches.

[0018] BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Embodiments of the invention will be better understood and readily apparent to one of ordinary skill in the art from the following written description, by way of example only, and in conjunction with the drawings, in which: Figure 1 shows a schematic diagram illustrating a “pipeline” for automated image annotation according to an example embodiment.

[0020] Figure 2 illustrates display of an image for user selection of screws on an eAxle, according to an example embodiment.

[0021] Figure 3A illustrates display of an auto-annotated image with screws on an eAxle identified, according to an example embodiment.

[0022] Figure 3B illustrates display of another auto-annotated image with screws on an eAxle identified, according to an example embodiment.

[0023] Figure 3C illustrates display of another auto-annotated image with screws on an eAxle identified, according to an example embodiment.

[0024] Figure 3D illustrates display of another auto-annotated image with screws on an eAxle identified, according to an example embodiment.

[0025] Figure 4 shows a schematic diagram illustrating the training phase, according to the example embodiment.

[0026] Figure 5 shows a schematic diagram illustrating the inference phase, according to the example embodiment.

[0027] Figure 6 shows a flowchart illustrating a method for automated image annotation, according to an example embodiment.

[0028] Figure 7 shows a schematic drawing illustrating a system for automated image annotation, according to an example embodiment.

[0029] DETAILED DESCRIPTION

[0030] Embodiments of the present invention provide a baseline method for significantly accelerating the labelling process and the application of deep neural networks.

[0031] In one embodiment, a method and system (also referred to as “pipeline” herein) for automated image annotation is provided. Suppose the user has a set of images S = {II, In} of images 100 to annotate, see Figure 1. With the method according to an example embodiment, the user labels one of the images 102 (that is representative of the rest of the data), and the pipeline tries to automatically infer labels for all of the other images, as indicated at numeral 104.

[0032] In particular, the example embodiment tackles the problem of object detection, meaning that a marker, e.g. a rectangle around each object of interest, is to be provided. On a high level, the method according to an example embodiment has two stages, that are referred to as “training” and “inference” herein for consistency with the Machine Learning terminology. During the training phase, the user is presented with one single image 102, and clicks on the objects to detect (e.g. the screws). Then, during the inference stage, the algorithm according to an example embodiment goes through all other images, and retrieves all image patches that are similar to any of the ones the user clicked, indicated at numeral 104.

[0033] Compared to the existing methods, example embodiments can have one or more of the following features / advantages:

[0034] (1) Working with any kind of objects: Unlike existing methods based on pre-trained neural networks, the method according to an example embodiment can work on objects that are not present in common datasets, like industrial parts. This is because conventional methods use a pretrained model to classify the user patches. The user clicks a few objects and they ask a neural network which classes these objects belong to. They then search for these categories in other images. It has been recognised by the present inventors that, when dealing with uncommon objects, like in industrial parts, the classes the neural networks provide are often wrong, and the neural network will likely predict the same category for any object that vaguely resembles the user input (since pretrained neural networks cannot differentiate screws and nuts for example), which leads to very low precision.

[0035] (2) Scalable: The method according to an example embodiment only requires the user to label a single image to train the algorithm. This is called “one-shot” learning. However, the method is scalable, meaning the user can label several images to boost the performances of the pipeline, in different embodiments.

[0036] (3) Multi-class labelling: The method according to an example embodiment can also natively handle several object categories. Thus, if the user clicks several objects and labels them as belonging to different categories (e.g. screw, nut, ...), the pipeline can retrieve each type of objects into the other images with little to no code modification.

[0037] (4) Applicable as-is: Finally, compared to geometry-based methods, the method according to an example embodiment does not require any prior knowledge about the objects to detect, and reduces the time dedicated to feature engineering.

[0038] Evaluation method of an example embodiment:

[0039] Since the method according to an example embodiment only relies on a single image to annotate all of the others, its accuracy will depend on the image used to train it. A proprietary dataset of 108 high-resolution eAxle images (1944x1200) was used in the evaluation. Note that this dataset has already been labelled in order to compare the method according to the example embodiments to a skilled human annotator. The objects to detect are screws located at various points around an e Axle cover. Each image of the dataset contains about 4 screws on average. It is noted that this task is especially hard, since screws are tiny objects (covering only a few pixels) and are difficult to spot as many parts of the eAxle have a similar color and geometry.

[0040] The dataset used was created by putting an eAxle on a turntable and taking images with various rotation angles (maximum 45°). For training, an image where the eAxle faces the camera (rotation close to 0°) was chosen as it is considered the most representative of the images in the dataset. It is noted that any of the images where the e Axle is centered and facing the camera was chosen, i.e. without evaluation / comparison of various images of that type within the dataset.

[0041] After training, the screws are inferred in the other images, and the Precision, Recall and Fl of the model is measured. The precision is the proportion of predictions that were actually correct, the recall is the ratio of screws that were able to be found and the total number of screws present in the images of the dataset, and Fl is the harmonic mean of precision and recall.

[0042] Quantitative Results for the example embodiment:

[0043] Using the evaluation method described above, the following result metrics was obtained:

[0044] These results for the example embodiment are very promising. Since there are about 4 screws per image on average, having a 83.8% precision means that our method predicts (1-0.838) * 4 = 0.65 false positives per image. In other terms, a significant proportion (35%) of the images will be perfectly annotated, and the rest will only require one click from the user to remove the false prediction. Accordingly, the method according to an example embodiment has the potential to save a tremendous amount of time for the user.

[0045] Qualitative Results for the example embodiment:

[0046] For the experiment using the example embodiment described above, the user provided 4 clicks on the single image 200 identifying the screws, as shown in Figure 2. However, it is noted that if there are more than 4 screws in the image, the user can benefit from clicking more / all of them, as more user inputs will benefit the model.

[0047] The method according to the example embodiment can either output segmentation masks (the precise pixels that it thinks belong to a screw) or rectangles around the screws (called bounding boxes), as a result of the inference. Figures 3A and B show the results on two images 300, 302 taken at different angles and with screws at various locations using segmentation masks, whereas Figures 3C and D show the results on two images 304, 306 taken at different angles and with screws at various locations using bounding boxes.

[0048] Training phase in more detail, according to the example embodiment:

[0049] Figure 4 illustrates the training phase, according to the example embodiment. SAM [9] was used, a very recent general-purpose image segmentation model, in order to segment the input image 400 into small patches, as indicated at numeral 402. In the meantime, the user clicks on the objects of interest in the image 400, as indicated at numeral 404. Then, all segments / patches on which the user has clicked are retrieved as indicated at numeral 406. Finally, a feature vector is computed for each of the clicked patches with a pretrained neural network 408. As for the model, DINOv2

[0010] was chosen for the example embodiment, noting that any recent architecture, like CLIP

[0011] or ConvNeXt V2

[0012] , can be used in different embodiments. As will be appreciated by a person skilled in the art, in deep learning, a feature vector (also called “embedding”) is the output of the neural network 408. The feature vector is a vector that represents the content of an image. When using a pretrained network 408, two similar patches should have similar feature vectors (low distance), and two patches that represent different objects should have different feature vectors (high distance). Hence, feature vectors enable to store a compact representation of the patches in the image, and to perform retrieval of patches with high accuracy. The feature vector e.g. 409 of each of the clicked patches e.g. 410 are stored in a database 412.

[0050] Inference phase according to the example embodiment:

[0051] With reference to Figure 5, for inference, the goal is to go through a set of new images 500, in this example embodiment all of the images in the dataset except for the one used for training, and retrieve patches that are similar to those the user clicked in the single training image (compare training phase description above). In the example embodiment, SAM [9], is used again to split all of the new images into patches, as indicated at numerals 502 and 504, and feature vectors 506 for all of those patches e.g. 508 are computed, using the neural network 408. It is noted that different models can be used for training and inference in different example embodiments. Then, the database 410 is searched for the nearest feature vector (e.g. in terms of cosine similarity in the example embodiment). If the distance between the current feature vector and the nearest one in the database 410 is less than a predefined threshold, then the patch is assumed to represent the same object, and e.g. a rectangle is drawn around it in the relevant image as indicated at numeral 510. If more than one type of object was trained to be inferred, the database 410 can be indexed according to the different object types and, for example, different colour and / or shapes can be used to indicate the inferred objects.

[0052] Figure 6 shows a flowchart 600 illustrating a method for automated image annotation, according to an example embodiment. At step 602, a training image is provided for user selection of one or more occurrences of at least one object within the image. At step 604, the training image is segmented into a first plurality of patches. At step 606, one or more patches out of the first plurality of patches are identified that correspond to respective ones of the user selected one or more occurrences of the at least one object. At step 608, a feature vector is computed, using a neural network, for each of the one or more patches out of the first plurality of patches. At step 610, occurrences of the at least one object in one or more sample images are inferred. At step 612, each of the one or more sample images is annotated based on each inferred occurrence of the at least one object for display; wherein inferring the occurrences of the at least one object comprises: segmenting each sample image into a second plurality of patches; compute, using the neural network, a feature vector for each of the patches of the second plurality of patches; and inferring the occurrences of the at least one object based on a comparison of the feature vectors for each of the patches of the second plurality of patches with the feature vectors for each of the one or more patches out of the first plurality of patches.

[0053] Inferring the occurrences of the at least one object may comprise computing respective distances between the feature vectors for each of the patches of the second plurality of patches and the feature vectors for each of the one or more patches out of the first plurality of patches. The occurrences may be inferred if one of the distances is less than a predefined threshold.

[0054] Annotating each of the one or more sample images may comprise generating a segmentation mask identifying the location of pixels belonging to each inferred occurrence of the at least one object.

[0055] Annotating each of the one or more sample images may comprise generating one or more markers identifying respective locations of each inferred occurrence of the at least one object.

[0056] Figure 7 shows a schematic drawing illustrating a system 700 for automated image annotation, according to an example embodiment. The system 700 comprises a display device 702 for providing a training image for user selection of one or more occurrences of at least one object within the image and for displaying annotated versions of one or more sample images; and a processing unit 704 configured to: segmenting the training image into a first plurality of patches; identifying one or more patches out of the first plurality of patches that correspond to respective ones of the user selected one or more occurrences of the at least one object; compute, using a neural network 706, a feature vector for each of the one or more patches out of the first plurality of patches; inferring occurrences of the at least one object in the one or more sample images; and annotating each of the one or more sample images based on each inferred occurrence of the at least one object for display of the annotated versions of the one or more sample images; wherein inferring the occurrences of the at least one object comprises: segmenting each sample image into a second plurality of patches; compute, using the neural network 706, a feature vector for each of the patches of the second plurality of patches; and inferring the occurrences of the at least one object based on a comparison of the feature vectors for each of the patches of the second plurality of patches with the feature vectors for each of the one or more patches out of the first plurality of patches.

[0057] The processing unit may be configured to infer the occurrences of the at least one object by computing respective distances between the feature vectors for each of the patches of the second plurality of patches and the feature vectors for each of the one or more patches out of the first plurality of patches. The processing unit may be configured to infer the occurrences if one of the distances is less than a predefined threshold.

[0058] The processing unit may be configured to annotate each of the one or more sample images by generating a segmentation mask identifying the location of pixels belonging to each inferred occurrence of the at least one object.

[0059] The processing unit may be configured to annotate each of the one or more sample images by generating one or more markers identifying respective locations of each inferred occurrence of the at least object.

[0060] Aspects of the systems and methods described herein, such image segmentation, feature extract! on / deep learning, image annotation and image display, may be implemented on computing device(s), including cloud-based computing device(s) and / or Internet-of-Things computing device(s), for example as functionality programmed into any of a variety of circuitry, including programmable logic devices (PLDs), such as field programmable gate arrays (FPGAs), programmable array logic (PAL) devices, electrically programmable logic and memory devices and standard cell-based devices, as well as application specific integrated circuits (ASICs). Some other possibilities for implementing aspects of the system include: microcontrollers with memory (such as electronically erasable programmable read only memory (EEPROM)), embedded microprocessors, firmware, software, etc. Furthermore, aspects of the system may be embodied in microprocessors having software-based circuit emulation, discrete logic (sequential and combinatorial), custom devices, fuzzy (neural) logic, quantum devices, and hybrids of any of the above device types. Of course the underlying device technologies may be provided in a variety of component types, e.g., metal-oxide semiconductor field-effect transistor (MOSFET) technologies like complementary metal-oxide semiconductor (CMOS), bipolar technologies like emitter-coupled logic (ECL), polymer technologies (e.g., silicon-conjugated polymer and metal-conjugated polymer-metal structures), mixed analog and digital, etc.

[0061] The various functions or processes disclosed herein may be described as data and / or instructions embodied in various computer-readable media, in terms of their behavioral, register transfer, logic component, transistor, layout geometries, and / or other characteristics. Computer-readable media in which such formatted data and / or instructions may be embodied include, but are not limited to, non-volatile storage media in various forms (e.g., optical, magnetic or semiconductor storage media) and carrier waves that may be used to transfer such formatted data and / or instructions through wireless, optical, or wired signaling media or any combination thereof. When received into any of a variety of circuitry (e.g. a computer), such data and / or instruction may be processed by a processing entity (e.g., one or more processors).

[0062] It will be appreciated by a person skilled in the art that numerous variations and / or modifications may be made to the present invention as shown in the specific embodiments without departing from the spirit or scope of the invention as broadly described. The present embodiments are, therefore, to be considered in all respects to be illustrative and not restrictive. Also, the invention includes any combination of features described for different embodiments, including in the summary section, even if the feature or combination of features is not explicitly specified in the claims or the detailed description of the present embodiments.

[0063] In general, in the following claims, the terms used should not be construed to limit the systems and methods to the specific embodiments disclosed in the specification and the claims, but should be construed to include all processing systems that operate under the claims. Accordingly, the systems and methods are not limited by the disclosure, but instead the scope of the systems and methods is to be determined entirely by the claims.

[0064] Unless the context clearly requires otherwise, throughout the description and the claims, the words "comprise," "comprising," and the like are to be construed in an inclusive sense as opposed to an exclusive or exhaustive sense; that is to say, in a sense of "including, but not limited to." Words using the singular or plural number also include the plural or singular number respectively. Additionally, the words "herein," "hereunder," "above," "below," and words of similar import refer to this application as a whole and not to any particular portions of this application. When the word "or" is used in reference to a list of two or more items, that word covers all of the following interpretations of the word: any of the items in the list, all of the items in the list and any combination of the items in the list.

[0065] References:

[0066] 1. Y. Lin, F. Lv, S. Zhu, M. Yang, T. Cour, K. Yu, L. Cao, T. S. Huang, “Large-scale image classification: Fast feature extraction and SVM training”, 2011.

[0067] 2. K. He, X. Zhang, S. Ren, J. Sun, “Deep Residual Learning for Image Recognition”, 2015.

[0068] 3. L. Fei-Fei, “ImageNet: crowdsourcing, benchmarking & other cool things”, 2010.

[0069] 4. V. Madam, “Introducing Amazon SageMaker Ground Truth”, 2018

[0070] 5. P. Tumas, A. Serackis, "Automated Image Annotation based on YOLOv3," 2018.

[0071] 6. H. Dong, X Chen, S. Sarkka, C. Stachniss, “Online pole segmentation on range images for long-term LiDAR localization in urban environments”, 2023

[0072] 7. E. Yildiz, F. Wbrgbtter, “DCNN-Based Screw Detection for Automated Disassembly Processes”, 2019.

[0073] 8. B. Missaoui, M. Noizet, P. Xu, “Map-aided annotation for pole base detection”, 2023.

[0074] 9. A. Kirillov et al., “Segment Anything”, 2023.

[0075] 10. M. Oquab et al., “DIN0v2: Learning Robust Visual Features without Supervision”, 2023. 11. A. Radford et al, “Learning Transferable Visual Models From Natural Language Supervision”, 2021.

[0076] 12. S. Woo, “ConvNeXt V2: Co-designing and Scaling ConvNets with Masked Autoencoders”, 2023.

Claims

CLAIMS1. A method for automated image annotation comprising the steps of: providing a training image (400) for user selection of one or more occurrences of at least one object within the image (400); segmenting the training image (400) into a first plurality of patches (402); identifying one or more patches (410) out of the first plurality of patches (402) that correspond to respective ones of the user selected one or more occurrences of the at least one object; computing, using a neural network, a feature vector (409) for each of the one or more patches (410) out of the first plurality of patches (402); inferring occurrences of the at least one object in one or more sample images (500); and annotating each of the one or more sample images (500) based on each inferred occurrence of the at least one object for display; wherein inferring the occurrences of the at least one object comprises: segmenting each sample image (500) into a second plurality of patches (504); compute, using the neural network, a feature vector (506) for each of the patches (508) of the second plurality of patches (504); and inferring the occurrences of the at least one object based on a comparison of the feature vectors (506) for each of the patches (508) of the second plurality of patches (504) with the feature vectors (409) for each of the one or more patches (410) out of the first plurality of patches (402).

2. The method of claim 1, wherein inferring the occurrences of the at least one object comprises computing respective distances between the feature vectors (506) for each of the patches (508) of the second plurality of patches (504) and the feature vectors (409) for each of the one or more patches (410) out of the first plurality of patches (402).

3. The method of claim 2, wherein the occurrences are inferred if one of the distances is less than a predefined threshold.

4. The method of any one of claims 1 to 3, wherein annotating each of the one or more sample images (500) comprises generating a segmentation mask (300, 302) identifying the location of pixels belonging to each inferred occurrence of the at least one object.

5. The method of any one of claims 1 to 3, wherein annotating each of the one or more sample images comprises generating one or more markers (304, 306) identifying respective locations of each inferred occurrence of the at least object.

6. A system (700) for automated image annotation comprising:a display device (702) for providing a training image (402) for user selection of one or more occurrences of at least one object within the image (402) and for displaying annotated versions (510) of one or more sample images (500); a processing unit (704) configured to: segmenting the training image (402) into a first plurality of patches (402); identifying one or more patches (410) out of the first plurality of patches (402) that correspond to respective ones of the user selected one or more occurrences of the at least one object; compute, using a neural network (706), a feature vector (409) for each of the one or more patches (410) out of the first plurality of patches (402); inferring occurrences of the at least one object in the one or more sample images (500); and annotating each of the one or more sample images (500) based on each inferred occurrence of the at least one object for display of the annotated versions (510) of the one or more sample images (500); wherein inferring the occurrences of the at least one object comprises: segmenting each sample image (500) into a second plurality of patches (504); compute, using the neural network, a feature vector (506) for each of the patches (508) of the second plurality of patches (504); and inferring the occurrences of the at least one object based on a comparison of the feature vectors (506) for each of the patches (508) of the second plurality of patches (504) with the feature vectors (409) for each of the one or more patches (410) out of the first plurality of patches (402).

7. The system of claim 6, wherein the processing unit (704) is configured to infer the occurrences of the at least one object by computing respective distances between the feature vectors (506) for each of the patches (508) of the second plurality of patches (504) and the feature vectors (409) for each of the one or more patches (410) out of the first plurality of patches (402).

8. The system of claim 7, wherein the processing unit (704) is configured to infer the occurrences if one of the distances is less than a predefined threshold.

9. The system of any one of claims 6 to 8, wherein the processing unit (704) is configured to annotate each of the one or more sample images (500) by generating a segmentation mask (300, 302) identifying the location of pixels belonging to each inferred occurrence of the at least one object.

10. The system of any one of claims 6 to 8, wherein the processing unit (704) is configured to annotate each of the one or more sample images (500) by generating one or more markers (304, 306) identifying respective locations of each inferred occurrence of the at least one object.

Citation Information

Patent Citations

  • Systems and methods for automatic image annotation

    US20230343438A1

  • Online, incremental real-time learning for tagging and labeling data streams for deep neural networks and neural network applications

    WO2018170512A1