Domain-specific human-model collaborative annotation tools
Through a personalized training process in the field of medical imaging, expert knowledge is transferred to inexperienced annotators, human experts and machine learning models are used for collaborative training, and attention maps and comparison functions are used to guide annotators, which solves the problem of shortage of experts in the annotation field and improves image labeling efficiency and annotation quality.
Patent Information
- Application Number
- CN201980101444.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-10-17
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2039-10-17
AI Technical Summary
In the field of medical imaging, the lack of experienced annotators leads to inefficient image labeling, and existing labeling tools cannot effectively solve the problem of shortage of experts in the annotation field.
Through a personalized training process, expert knowledge is transferred to inexperienced annotators, human experts and machine learning models are used for collaborative training, attention maps and comparison functions are used to guide annotators, personalized feedback and evaluation are provided, and annotation efficiency is improved.
It improves the efficiency of the image labeling process, alleviates the shortage of experts in the annotation field, enhances the quality of annotation, and improves the training effect of machine learning models.
Smart Images

Figure CN114600196B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to annotation tools, and in particular to domain-specific human-model co-annotation tools that train human annotators and machine learning systems to improve the efficiency of the domain-specific image labeling process. Background Art
[0002] In the field of medical imaging, deep learning is widely used to solve classification, detection, and segmentation problems. Labeled (annotated) data is crucial for training deep learning models. However, the type of medical image data varies depending on the type of imaging device used and the anatomy / tissue being examined, which increases the difficulty of labeling such domain-specific data.
[0003] Domain-specific image annotation requires specialized training and domain knowledge from annotators. The annotator's experience level significantly impacts annotation quality. Unfortunately, the lack of experienced annotators to label diverse biomedical data poses challenges in providing efficient assessment and treatment.
[0004] Currently, there are several general-purpose labeling tools for labeling images. One set of labeling tools uses hand-drawn annotations. For example, the LabelIMG tool supports bounding boxes and one-class labels. The VGG image annotator has options for adding object and image attributes or labels. Other labeling tools, such as Supervise.ly and Labelbox, use models to provide semantic segmentation and help predict labels for model training through human verification. Other labeling tools use active learning or reinforcement learning to train models using a small number of labeled images. Active learning models select uncertain examples and seek help from human reviewers to complete the labeling. To produce more accurate predictions, machine learning models are used in systems such as AWS SageMaker Ground Truth and Huawei Cloud ModelArts. These systems provide annotation tools that select images for display to human annotators and use the newly labeled images to further train the machine learning model. The Polygon RNN++ segmentation tool sequentially generates the vertices of polygons to outline objects. This allows human annotators to intervene over time and correct the vertices as needed to produce more accurate segmentations.
[0005] Unfortunately, such state-of-the-art systems generally do not help address the shortage of annotation domain experts. Summary of the Invention
[0006] Various examples are now described to introduce some concepts in a simplified form that will be further described in the detailed description. This summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter.
[0007] An end-to-end domain-specific human model co-annotation system and method are described. This system and method trains human annotators to improve the efficiency of the domain-specific image labeling process, thereby addressing the shortage of human domain experts. The human model co-annotation system described herein transfers expert knowledge to new human annotators through a personalized training process, while providing additional samples for training machine learning systems. In an exemplary embodiment, the human model co-annotation system includes at least the following features:
[0008] 1. An annotation system comprising an annotator training system that provides image samples for inexperienced annotators to label. The sampling and training process is personalized and based on the annotators' errors. Attention maps are used to guide the annotators to learn from their errors.
[0009] 2. The evaluation phase in the annotator training system scores the annotators and provides reliable personalized feedback to the annotators based on the annotations they submit.
[0010] 3. By integrating learned domain knowledge and general human intelligence, newly trained annotators can add labels to help expand the pool of labeled samples and improve the annotation system training model.
[0011] According to a first aspect of the present invention, a training method for training a human annotator to annotate an image is provided. The training method comprises: presenting an image sample to the human annotator for annotation, wherein the image sample has been previously annotated by at least one of an expert human annotator and a machine learning annotator; receiving one or more suggested annotations from the human annotator; comparing the one or more suggested annotations of the human annotator with previous annotations of the image sample by the expert human annotator or the machine learning annotator; presenting an attention map to draw the human annotator's attention to annotation errors identified by the comparison; and selecting a next image sample based on any errors identified in the comparison.
[0012] According to a second aspect of the present invention, a human-model collaborative annotation system is provided, comprising: a database storing images previously annotated by at least one of an expert human annotator and a machine learning annotator; a display displaying an image selected from the database; an annotation system for enabling a human annotator to annotate an image presented on the display; and an annotation training system. The annotation training system: selects an image sample from the database to be displayed on the display for annotation by a human annotator; receives one or more suggested annotations from the annotation system; compares the one or more suggested annotations of the human annotator with previous annotations of the image sample by an expert human annotator or a machine learning annotator; presents an attention map on the display to alert the human annotator to any annotation errors identified by the comparison; and selects a next image sample from the data block based on any errors identified in the comparison.
[0013] According to a third aspect of the present invention, there is provided a non-transitory computer-readable medium storing computer instructions for training a human annotator to annotate images, which, when executed by one or more processors, causes the one or more processors to perform the following operations:
[0014] Presenting an image sample to a human annotator for annotation, wherein the image sample has been previously annotated by at least one of an expert human annotator and a machine learning annotator; receiving one or more suggested annotations from the human annotator; comparing the one or more suggested annotations of the human annotator with previous annotations of the image sample by the expert human annotator or the machine learning annotator; presenting an attention map to draw the attention of the human annotator to annotation errors identified by the comparison; and selecting a next image sample based on any errors identified in the comparison.
[0015] In a first implementation of any of the above aspects, annotation performance of a human annotator is evaluated by applying a weighting function and a numerical metric to comparison results of one or more suggested annotations by the human annotator with previous annotations of an image sample by an expert human annotator or a machine learning annotator.
[0016] In a second implementation of any of the above aspects, when a human annotator is evaluated as having annotation performance above a threshold, image samples for annotation by the human annotator are presented, and the annotated image samples from the human annotator are added to an image sample pool, the image sample pool including image samples previously annotated by expert human annotators or machine learning annotators.
[0017] In a third implementation of any of the above aspects, the annotated image samples from the human annotators added to the pool include weights based on the annotation performance of the human annotators.
[0018] In a fourth implementation of any of the above aspects, the human annotator is certified for future annotation tasks when the human annotator's annotation performance is above a predetermined level for the type of annotation the human annotator has been trained on.
[0019] In a fifth implementation of any of the above aspects, the annotation performance of multiple human annotators is compared on the same set of images to establish a quality metric for the multiple human annotators.
[0020] In a sixth implementation of any of the above aspects, the attention map is presented to a display to provide personalized explanations of annotation errors.
[0021] In a seventh implementation of any one of the above aspects, the image to be annotated includes a medical image, a geographic image, and / or an industry image.
[0022] The method can be executed and the instructions in the computer-readable medium can be processed by the system to train an annotator, such as a medical imaging annotator, and further features of the method and the instructions in the computer-readable medium are generated by the function of the system. In addition, the explanations provided for each aspect and its implementation are also applicable to the other aspects and corresponding implementations. Different embodiments can be implemented in hardware, software, or any combination thereof. In addition, any of the above examples can be combined with any one or more of the other examples to create new embodiments within the scope of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In the drawings, which are not necessarily drawn to scale, like numerals may describe similar components in different views.The drawings illustrate generally, by way of example, and not by way of limitation, various embodiments described herein.
[0024] Figures 1A to 1C shows images showing different types of image annotations, including classification ( Figure 1A )、Detection( Figure 1B ) and split( Figure 1C ).
[0025] Figure 2 A block diagram of an exemplary embodiment of a human model collaborative annotation system is shown.
[0026] Figure 3A A flow chart of a method for generating annotated images for annotator training in an exemplary embodiment is shown.
[0027] Figure 3B A flow chart illustrating the operation of an annotator training system in an exemplary embodiment is shown.
[0028] Figures 4A to 4C Unannotated sample images are shown, including normal lung images ( Figure 4A), lung images with bilateral pleural effusion ( Figure 4B ), and lung images with opacities in the lungs ( Figure 4C ).
[0029] Figure 4D A segmented lung image is shown.
[0030] Figure 4E Lung images are shown with boxes showing the ground truth annotations used by domain experts for disease detection in the samples.
[0031] Figures 4F to 4G shows machine-generated attention maps that show what the human annotators missed during annotation, including bilateral pleural effusions ( Figure 4F ) and lung opacities ( Figure 4G ).
[0032] Figure 5 A block diagram of a circuit for performing a method according to an exemplary embodiment is shown. DETAILED DESCRIPTION
[0033] First, it should be understood that although the following provides an illustrative implementation of one or more embodiments, the present invention disclosed in FIG. Figure 5 The described systems and / or methods may be implemented using any number of techniques, whether currently known or existing. The present invention should not be limited to the illustrative implementations, drawings, and techniques described below, including the exemplary designs and implementations illustrated and described herein, and may be modified within the scope of the appended claims along with their full scope of equivalents.
[0034] In one embodiment, the functions or algorithms described herein can be implemented in software. The software may include computer-executable instructions stored in a computer-readable medium or a computer-readable storage device (such as one or more non-transitory memories or other types of hardware-based local or network storage devices). In addition, these functions may correspond to modules, which may be software, hardware, firmware, or any combination thereof. As needed, multiple functions may be performed in one or more modules, and the described embodiments are merely examples. The software may be executed on a digital signal processor, an ASIC, a microprocessor, or other types of processors running on a computer system (such as a personal computer, server, or other computer system), transforming such a computer system into a specifically programmed machine.
[0035] Traditional annotation tools of the type described above typically limit the interaction between humans and annotation models to labeling (annotating) images. For some tools, only a simple explanation of the labels is provided. For domain-specific labeling tasks, traditional annotation tools, such as those mentioned above, do not help eliminate the obstacles for annotators (especially inexperienced annotators). Such systems do not transfer knowledge and do not provide guidance or evaluation capabilities. In general, traditional annotation tools cannot address the shortage of annotation domain experts by training human annotators or improving machine learning models.
[0036] The human-model co-annotation system described in this paper uses human experts and a machine learning model to train non-expert human annotators, while the human experts simultaneously train the machine learning model. An automated guidance system uses comparison functions and attention maps to guide the non-expert annotators, alerting them to any annotation errors. The guidance system is personalized for the annotators and includes interactive learning elements, including a comparison function that lists and compares positive and negative examples, or examples with different labels, allowing the annotator to identify the differences. The guidance system also leverages attention maps from the trained model to highlight signals for a given label to guide the annotator. Annotators can zoom in and out and select examples to review based on where they made errors, allowing them to get up to speed faster. The guidance system also includes a comprehensive annotator evaluation system that guides training, certifies annotators, and weights their contributions to the annotated image database after training. The guidance system further explains the reasoning behind the labels and attention maps to enhance training.
[0037] In an exemplary embodiment, the annotation system provides services including a ground truth collection process, wherein, for a given biomedical image labeling task, a machine learning model can be trained based on samples labeled by experts with domain knowledge. The ground truth collection process is supplemented by a personalized guidance process that timely guides inexperienced annotators who lack domain knowledge to understand how to label images, thereby transferring the expert's knowledge to new annotators. An evaluation of the annotator's performance is also provided to generate an objective evaluation score. The evaluation score can then be used to further train new annotators and further refine the machine learning model using new labeled data constrained by weighting factors. The training system determines the next sample to be displayed based on the annotator's error and highlights the annotation errors using an attention map. For example, the next sample can be an image of the same type as a similar attention map. The samples will have the same type of attention region used to annotate the image.
[0038] The image annotations described in this article can take different forms and be applied to different types of images. Figures 1A to 1C Three main types of image annotation are shown, including classification ( Figure 1A)、Detection( Figure 1B ) and split( Figure 1C ). Image level annotation is used for classification. In this type of annotation, an image is given a label (e.g. brain tumor) if the label is included in the image’s content. For example, Figure 1A The categories of "outdoors", "horses", "grass" and "people" are shown for images containing each of these elements. Figure 1B The bounding box annotation shown can be used for detection. In this type of annotation, a rectangle is drawn to tightly surround an object in the image. The rectangle typically contains the object of interest, such as "person," "horse," or "tumor region." Figure 1C Segmented contour annotation is shown. In this type of annotation, a polygon is drawn around the outline of an object to delineate the object in the image corresponding to a label. The polygon typically contains a region of interest, such as "person," "horse," or "tumor." In each case, the object is matched to the label. The systems and methods described herein support all three types of annotation.
[0039] Figure 2 A block diagram of an exemplary embodiment of a human model collaborative annotation system 200 is shown. In the system 200, unlabeled image data (D1) from a database 210 is provided to an expert for labeling using an expert computer system having annotation software 220. The resulting images labeled by the expert (A1) are provided to a labeled image database 230 to create a database of annotated images. The labeled images are also used to train a machine learning model of a machine learning system 240, which, after being trained, can in turn receive unlabeled images from the database 210 and generate more annotated images to be stored in the labeled image database 230. Furthermore, after the machine learning model is trained, the collaboratively trained machine learning model of the machine learning system 240 can be deployed for use by non-expert annotators.
[0040] In an exemplary embodiment, system 200 also includes an annotation system 250 that selects samples to present to non-expert human annotators on display 260 for labeling. In an exemplary embodiment, non-expert human annotators can receive feedback from annotator training system 270, which compares annotated images generated by the non-expert human annotators using annotation system 250 with corresponding expert-annotated images provided by labeled image database 230. The feedback includes attention maps that guide the non-expert human annotators in learning from their mistakes. The evaluator's performance is evaluated, and the next image to be annotated by the non-expert human annotator on annotation system 250 is selected based on the non-expert human annotator's performance in labeling the previously presented image. When a non-expert human annotator receives consistently good ratings from the performance evaluation system, the non-expert human annotator can be added to the pool of expert annotators and supported in annotating additional images that may be added to the pool of labeled samples in labeled image database 230. The added images can be weighted based on the accuracy of the non-expert human annotators' annotations, thereby improving the trained model.
[0041] In an exemplary embodiment, annotations for classification can be captured by the annotator entering a label in a text box or checking a yes or no button for a given label. Annotations for classification rectangles can be captured by the annotator dragging the mouse from the top left corner to the bottom right corner. Annotations for segmentation can be captured by clicking on the vertices of a polygon. Comparison can be performed by comparing the annotator's suggested label Z with the ground truth label G, where G and Z are labels for classification, rectangles for detection, and contours / polygons for segmentation.
[0042] The performance score is calculated by comparing the ground truth label G with the annotator label Z and averaging it over all images annotated by the annotator. For classification, it is correct only if Z is equal to G. For detection and segmentation, the intersection over union (IOU) is used, which is calculated by dividing the area of the intersection of Z and G by the area of the union of Z and G. The annotation is correct only if the IOU is greater than a predefined threshold (such as 0.5). The performance score of the annotator is then used to weight the annotator's annotations if it is greater than a predefined threshold (such as 99%).
[0043] Figure 3A 1 is a flow chart of a method 300 for generating annotated images for annotator training in an exemplary embodiment. The method 300 starts at 310 and receives an unlabeled image set D from the unlabeled image database 210 at 320. u At 330, the uThe sampled subset is represented as Dl, D u Update by deleting D1. At 340, the sample subset D1 is provided to the annotation system 220 of the domain expert annotator so that the expert annotator can annotate the samples in D1. The resulting annotations A1 are provided to the labeled image database 230. At 350, the expert annotations A1 of the sample images in D1 can also be used to train the machine learning model of the machine learning system 240. At 360, the machine learning (ML) model can also be used to train the machine learning model in D1. u More samples are labeled in D, but no samples are labeled in Dl. D generated by the machine learning model u The annotated samples in can be added to the labeled image database 230 along with the annotations labeled by the machine learning model. Then, at 370, the labeled images in the labeled image database 230 are ready for use by the annotator training system 270. Figure 3B The operation of the annotator training system 270 is described.
[0044] As described above, the annotator training system 270 teaches new annotators the domain knowledge of the expert annotators. During the first iteration of the guidance, the non-expert human annotators are provided with randomly selected samples from the labeled image database 230 to label, and the automatic evaluation system evaluates the performance of the non-expert human annotators as described above. Based on the performance, the subsequent guidance process is personalized as follows:
[0045] 1. Personalized Sampling: More images are sampled that include the same types of features that led to errors for non-expert human annotators in previous rounds of annotation and that may demonstrate some confusion. Unlike the active learning paradigm used in traditional annotation systems (which simply involves having the model select the most uncertain samples for the annotator to label), the personalized annotation training system described in this paper addresses the confusion or uncertainty exhibited by specific non-expert human annotators during training.
[0046] 2. Personalized evaluation: The evaluation function compares the annotator’s errors with the known labels and uses attention maps (as shown in Figure 4) to highlight differences and areas for the annotator to learn from (using general human cognitive intelligence).
[0047] 3. Personalized Guidance: The system differentiates the learning process by selecting more samples based on each annotator’s own cognitive intelligence and learning speed. This makes learning more effective and engaging, thereby training annotators faster and making the annotation process more efficient.
[0048] Figure 3BA flowchart illustrating the operation of the annotator training system 270 in an exemplary embodiment is shown. As shown, the method begins at 372 by extracting a random sample of images X and their annotations A from the labeled image database 230. At 374, X and A are presented to a non-expert human annotator, where the annotations are highlighted in the original images as an attention map. At 376, another random sample of images X and their annotations A is extracted. Only images X are presented to the non-expert human annotator, who is asked to label the images with annotations Y. At 378, annotations Y are evaluated to assess the performance of the non-expert human annotator by comparing the annotator's annotations Y with the expert or machine learning model annotations A, as described above. Optionally, at 378, annotations Y can be compared with annotations from other annotators to evaluate the annotation performance of multiple human annotators on the same set of images to establish a quality metric for the multiple human annotators.
[0049] It should be understood that evaluation is a key component of the guidance system, which should comprehensively analyze the performance of the annotators. In exemplary embodiments, evaluation metrics may include, but are not limited to, the accuracy of the annotator on each task and each label (e.g., detection and segmentation, and Intersection over Union (IOU) which measures the degree of overlap in the image), the overall quality of the annotator's labels, and the learning curve of the annotator (including the history of speed and accuracy).
[0050] After the performance of the annotator is evaluated at 378, a check is made at 380 of the method to determine if the performance of the annotator is above a predefined threshold. If the performance of the annotator is above the predefined threshold, then at 382 the annotator can be added to the expert domain ( Figure 2 At 220 ). However, the annotator is also assigned a rank level that can be used to weight its future annotations. For example, based on the annotator's performance during annotation training, future annotations can be adjusted by a weight between 0 and 1. Optionally, at 384 , the annotator training system 270 can certify the annotator according to a predetermined certification process designed to provide an objective measure of "expert" annotators. The annotator training process then ends at 386 .
[0051] On the other hand, at 380, when the annotator's performance is below a predefined threshold, training can continue. Further training includes, at 388, generating an attention map that highlights annotation errors by comparing the annotator's annotations Y and the expert or machine learning model annotations A. At 390, the attention map is presented to the annotator, where errors and explanations for those errors are highlighted to provide personalized training. Then, at 392, the annotator training system 270 determines whether to continue training. If the annotator's performance is not above the predefined threshold, then at 386, training ends. However, if training is to continue, the process will return to 376 to extract another random sample of another batch of images and their annotations to repeat the process.
[0052] In an exemplary embodiment, the annotation training system 270 also provides guidance on what sample images to show next based on errors highlighted in one or more previous iterations of the training process. For example, in the case of image classification, the annotation training system 270 may present an annotator with image X, where the ground truth label for image X is known to be Y but not displayed to the annotator. The annotator mistakenly assigns a label Z that is different from Y. The system generates an attention map in which regions associated with Y are highlighted in X, and displays image X and the highlighted attention map to guide the annotator. The annotation training system 270 then provides a sample of another image with the same label Y to reinforce the training for errors in identifying label Y. For detection and segmentation, the process is similar, except that the label Y is presented as a rectangle for detection and a contour for segmentation. This process further enables personalized training by presenting images that highlight areas where the annotator encountered difficulties. Through this evaluation system, the annotation training system 270 learns how to train the annotator in a personalized manner in the next stage and how to integrate the annotator's labels with other labels under the constraints of weighting factors. The annotation training system 270 may also certify annotators for future tasks requiring the same expertise.
[0053] Figures 4A to 4C Unannotated sample images are shown, including normal lung images ( Figure 4A ), lung images with bilateral pleural effusion ( Figure 4B ), and lung images with opacities in the lungs ( Figure 4C ). Figure 4D A segmented lung image is shown at 400 . Figure 4E Lung images are shown with boxes showing the ground truth annotations used by domain experts for disease detection in the samples. Figures 4F to 4G shows machine-generated attention maps that show what the human annotators missed during annotation, including bilateral pleural effusions ( Figure 4F ) and lung opacities ( Figure 4G ). This attention map allows the annotator to focus on areas of interest, focusing on errors made by the annotator. In an exemplary embodiment, for improved training, the attention map can also be presented with explanations of correct annotations and / or incorrect annotations.
[0054] In an exemplary embodiment, the attention map uses different colors or intensity values to highlight areas at 410 to guide the user to the correct area to correct the annotation. The attention map may include the following features:
[0055] The image is given a label due to the highlighted areas;
[0056] The annotator made an error in the highlighted area.
[0057] Many methods are used in the art to obtain these attention maps. For example, given an input image to a well-trained machine learning model (such as a deep neural network), the machine learning model can output a label with a ground truth that exactly matches the annotations of the domain expert. This process is called the forward pass from the input image to the output prediction. By replacing the output prediction with the ground truth and performing the process in reverse, the input image can be recovered, where the areas related to the ground truth are highlighted. Therefore, the machine-generated attention map is used to show the human annotator what was missed during the annotation process.
[0058] The guidance component and human model collaborative annotation system described in this article provide several advantages over traditional annotation systems. For example, the human model collaborative annotation system described in this article can alleviate or solve the shortage of expert annotators. The human model collaborative annotation system also eliminates or reduces barriers to domain-specific annotation tasks (including annotation of medical images, geographic images, and industry images) through knowledge transfer between machine intelligence and human intelligence (expert intelligence and general human cognitive intelligence). In addition, trained annotators can obtain certification to perform future annotation tasks that require the same or similar expertise. Even without full certification, trained annotators can use weighted annotations to create new training images for use by the training system.
[0059] The evaluation system can also be used to assess the labeling quality of multiple workers. By comparing annotations with each other, the human-model co-annotation system can help integrate labels from different workers at different times and locations.
[0060] It will also be understood that the human-model collaborative annotation system can utilize machine learning models and deep learning algorithms to transfer domain expert knowledge to others. In addition to being used for the annotation described herein, the human-model collaborative annotation system described herein can also be used for education and industrial professional training.
[0061] Figure 5 A block diagram of an exemplary machine 500, such as an annotation training system, is shown, on which any one or more of the techniques (e.g., methods) discussed herein can be executed. In alternative embodiments, machine 500 can operate as a standalone device or can be connected (e.g., networked) to other machines. In a network deployment, machine 500 can operate in the capacity of a server machine, a client machine, or both in a server-client network environment. In one example, machine 500 can act as a peer machine in a peer-to-peer (P2P) (or other distributed) network environment. Machine 500 can be a personal computer (PC), a tablet PC, a set-top box (STB), a personal digital assistant (PDA), a mobile phone, a network appliance, an IoT device, an automotive system, or any machine capable of executing instructions (sequential or otherwise) specifying actions to be taken by the machine. In addition, although only a single machine is shown, the term "machine" should also be construed to include any collection of machines that individually or jointly execute a set (or multiple sets) of instructions to perform any one or more of the methods discussed herein, such as cloud computing, software as a service (SaaS), or other computer cluster configurations.
[0062] The examples described herein may include, or may be operated by, logic, components, devices, packages, or mechanisms. A circuit is a collection (e.g., set) of circuits implemented in a tangible entity, including hardware (e.g., simple circuits, gates, logic, etc.). Circuit membership may change over time and with potential hardware variability. A circuit includes components that can perform specific tasks individually or in combination when operated. In one example, the hardware of a circuit can be immutably designed to perform a specific operation (e.g., hardwired). In one example, the hardware of a circuit can include variably connected physical components (e.g., execution units, transistors, simple circuits, etc.) that include computer-readable media with physical modifications (e.g., magnetic, electrical, movable placement of constant mass particles, etc.) to encode instructions for specific operations. When the physical components are connected, the underlying electrical properties of the hardware components are changed, for example, from an insulator to a conductor, or from a conductor to an insulator. The instructions enable the participating hardware (e.g., execution units or loading mechanisms) to create components of the circuit in hardware through variable connections to perform portions of a specific task when operated. Thus, when the device is operating, the computer-readable medium is communicatively coupled to other components of the circuit. In one example, any physical component can be used in multiple members of multiple circuits. For example, during operation, an execution unit can be used in a first circuit in a first circuit system at one point in time and reused by a second circuit in the first circuit system, or reused by a third circuit in the second circuit system at a different time.
[0063] The machine (e.g., computer system) 500 (e.g., the annotation system 250, the annotator training system 270, etc.) may include a hardware processor 502 (e.g., a CPU, a graphics processing unit (GPU), a hardware processor core, or any combination thereof), a main memory 504, and a static memory 506, some or all of which may communicate with each other via an interconnection link (e.g., a bus) 508. The machine 500 may also include a display device 510, an alphanumeric input device 512 (e.g., a keyboard), and a user interface (UI) navigation device 514 (e.g., a mouse). In one example, the display unit 510, the input device 512, and the UI navigation device 514 may be a touch screen display. The machine 500 may also include a signal generating device 518 (e.g., a speaker), a network interface device 520, and one or more sensors 516, such as a global positioning system (GPS) sensor, a compass, an accelerometer, or other sensors. The machine 500 may include an output controller 528, such as a serial (e.g., USB) connection, a parallel connection, or other wired or wireless (e.g., infrared (IR), near field communication (NFC), etc.) connection, for communicating with or controlling one or more peripheral devices (e.g., a printer, a card reader, etc.).
[0064] The machine 500 may include a machine-readable medium 522 having stored therein one or more sets of data structures or instructions 524 (e.g., software) that embody or are used by any one or more of the techniques or functionality described herein. The instructions 524 may also reside, completely or at least partially, within the main memory 504, within the static memory 506, or within the hardware processor 502 during execution by the machine 500. In one example, one or any combination of the hardware processor 502, the main memory 504, or the static memory 506 may constitute the machine-readable medium 522.
[0065] Although machine-readable medium 522 is illustrated as a single medium, the term “machine-readable medium” may include a single medium or multiple media (eg, a centralized or distributed database, or associated caches and servers) for storing one or more instructions 524 .
[0066] The term "machine-readable medium" may include any medium capable of storing or encoding instructions for execution by the machine 500 and causing the machine 500 to perform any one or more of the techniques of the present invention, or any medium capable of storing, encoding, or carrying data structures used by or related to such instructions. Non-limiting examples of machine-readable media include solid-state memory, optical, and magnetic media. In one example, a bulk machine-readable medium includes a machine-readable medium having a plurality of particles with a constant (e.g., stationary) mass. Thus, a bulk machine-readable medium is not a temporary storage of propagating signals. Specific examples of bulk machine-readable media may include: non-volatile memory, such as semiconductor memory devices (e.g., electrically programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM)), and flash memory devices; magnetic disks, such as internal hard disks and removable disks; magneto-optical disks; CD-ROM and DVD-ROM disks.
[0067] Instructions 524 (e.g., software, programs, operating systems (OS), etc.) or other data are stored in storage device 521 and can be accessed by memory 504 for use by processor 502. Memory 504 (e.g., DRAM) is typically fast but volatile, and therefore a different type of storage than storage device 521 (e.g., SSD), which is suitable for long-term storage and can be used even when powered off. Instructions 524 or data used by a user or machine 500 are typically loaded into memory 504 for use by processor 502. When memory 504 is full, virtual space from storage device 521 can be allocated to supplement memory 504. However, because storage device 521 is typically slower than memory 504, with write speeds typically at least twice as fast as read speeds, there is storage device latency (compared to memory 504 (e.g., DRAM)). Therefore, the use of virtual memory can significantly reduce the user experience. Furthermore, using storage device 521 for virtual memory can significantly shorten the useful life of storage device 521.
[0068] Compared to virtual memory, virtual memory compression (e.g. A kernel feature called "ZRAM" stores some memory as compressed blocks to avoid paging to storage device 521. Before writing this data to storage device 521, paging occurs in compressed blocks. Virtual memory compression increases the available size of memory 504 while reducing wear on storage device 521.
[0069] The instructions 524 may also be sent or received over a communication network 526 using a transmission medium via a network interface device 520 utilizing any of a number of transmission protocols (e.g., frame relay, internet protocol (IP), transmission control protocol (TCP), user datagram protocol (UDP), hypertext transfer protocol (HTTP), etc.). Exemplary communication networks may include a local area network (LAN), a wide area network (WAN), a packet data network (e.g., the Internet), a mobile phone network (e.g., a cellular network), a plain old telephone service (POTS) network, and a wireless data network (e.g., the Institute of Electrical and Electronics Engineers (IEEE) 802.11 series of standards known as IEEE 802.16 series of standards is called ), IEEE 802.15.4 series standards, P2P networks, etc. In one example, the network interface device 520 may include one or more physical jacks (e.g., Ethernet, coaxial, or telephone jacks) or one or more antennas to connect to the communication network 526. In one example, the network interface device 520 may include multiple antennas to perform wireless communications using at least one of single-input multiple-output (SIMO), multiple-input multiple-output (MIMO), or multiple-input single-output (MISO) technology.
[0070] The term "transmission medium" shall be taken to include any intangible medium that can store, encode, or carry instructions for execution by the machine 500, and includes digital or analog communications signals or other intangible medium to facilitate communication of such software.
[0071] The above detailed description includes reference to the accompanying drawings, which form a part of the detailed description. The accompanying drawings show, by way of illustration, specific embodiments in which the systems and methods described herein may be implemented. These embodiments are also referred to herein as "examples." In addition to the elements shown or described, these examples may also include other elements. However, the present inventors have also contemplated examples in which only those elements shown or described are provided. In addition, with respect to a particular example (or one or more aspects thereof), or with respect to other examples (or one or more aspects thereof) shown or described herein, the present inventors have also contemplated examples in which any combination or arrangement of those elements (or one or more aspects thereof) is used.
[0072] Herein, as is common in patent documents, the terms "a" or "an" are used to include one or more, independent of any other instance or usage of "at least one" or "one or more." Herein, the term "or" is used to refer to a non-exclusive or, thus, unless otherwise stated, "A or B" may mean "including A but not B," "including B but not A," and "including A and B." In the appended claims, the terms "including" and "in which" are used as synonyms for the respective terms "comprising" and "wherein." Furthermore, in the appended claims, the terms "including" and "comprising" are open-ended, i.e., systems, devices, articles, or processes that include elements in addition to those listed after those terms in the claim are still considered to fall within the scope of the claim. Furthermore, in the appended claims, the terms "first," "second," and "third," etc. are used merely as labels and are not intended to impose numerical requirements on their objects.
[0073] In various examples, the components, controllers, processors, units, engines, or tables described herein may include physical circuitry or firmware stored in a physical device. As used herein, a "processor" refers to any type of computing circuit, such as, but not limited to, a microprocessor, a microcontroller, a graphics processor, a digital signal processor (DSP), or any other type of processor or processing circuit, including a group of processors or a multi-core device.
[0074] It should be understood that when an element is referred to as being "on," "connected to," or "coupled to" another element, the element can be directly on, connected to, or coupled to the other element, or intervening elements may be present. In contrast, when an element is referred to as being "directly on," "directly connected to," or "directly coupled to" another element, there are no intervening elements or layers present. If two elements are shown as connected by a line in a drawing, the two elements may be coupled or directly coupled unless otherwise specified.
[0075] The method examples described herein may be at least partially machine or computer implemented. Some examples may include a computer-readable medium or machine-readable medium encoded with operable instructions to enable an electronic device to perform the methods described in the above examples. The implementation of these methods may include code, such as microcode, assembly language code, high-level language code, etc. These codes may include computer-readable instructions for performing various methods. The code may form part of a computer program product. In addition, for example, during execution or at other times, the code may be tangibly stored in one or more volatile or non-volatile tangible computer-readable media. Examples of these tangible computer-readable media may include, but are not limited to, hard disks, removable disks, removable optical disks (such as CDs and DVDs), cassettes, memory cards or sticks, RAM, ROM, SSDs, UFS devices, eMMC devices, etc.
[0076] The above description is intended to be illustrative, not restrictive. For example, the above embodiments (or one or more aspects thereof) may be used in combination with each other. For example, after reading the above description, one of ordinary skill in the art may use other embodiments. It is submitted with the understanding that it will not be used to interpret or limit the scope or meaning of the claims. In addition, in the above detailed description, various features may be grouped together to simplify the invention. This should not be interpreted as meaning that unclaimed disclosed features are essential to any claim. On the contrary, the subject matter of the invention may be less than all the features of a particular disclosed embodiment. Therefore, the appended claims are hereby incorporated into the detailed description, each claim existing separately as a separate embodiment, and it is conceivable that these embodiments may be combined with each other in various combinations or permutations. The scope of the invention should be determined with reference to the appended claims and the full scope of equivalents to these claims.
Claims
1. A training method for training human annotators to annotate images, characterized in that include: presenting an image sample to a human annotator for annotation, wherein the image sample has been previously annotated by at least one of an expert human annotator and a machine learning annotator; receiving one or more suggested annotations from the human annotator; comparing the one or more suggested annotations of the human annotator with previous annotations of the image sample by the expert human annotator or the machine learning annotator; evaluating the annotation performance of the human annotator using a weighting function and a numerical metric based on a result of the comparison; If the annotation performance is below a preset threshold, presenting an attention map to highlight attention areas to draw the attention of the human annotator to annotation errors identified by the comparison; wherein the annotating the errors comprises providing a personalized explanation of the annotation errors on a display using the attention map; A next image sample is selected based on any errors identified in the comparison, the next image sample having the same type of attention region as the attention map for reinforcement training for annotation errors identified by the comparison.
2. The method according to claim 1, characterized in that Also included is presenting image samples for annotation by the human annotator when the human annotator is evaluated as having annotation performance above a threshold, and adding the annotated image samples from the human annotator to an image sample pool, the image sample pool including image samples previously annotated by the expert human annotator or the machine learning annotator.
3. The method according to claim 2, characterized in that The annotated image samples from the human annotator added to the pool of image samples include weights based on the annotation performance of the human annotator.
4. The method according to claim 1, wherein Also included is certifying the human annotator for future annotation tasks when the human annotator's annotation performance is above a predetermined level for the type of annotation for which the human annotator has been trained.
5. The method according to claim 1, wherein Also included is comparing annotation performance of multiple human annotators for the same set of images to establish a quality metric for the multiple human annotators.
6. The method according to claim 1, wherein The image to be annotated includes at least one of a medical image, a geographical image, and an industry image.
7. A human-model collaborative annotation system, characterized in that include: a database storing images previously annotated by at least one of an expert human annotator and a machine learning annotator; a display that displays an image selected from the database; an annotation system for enabling a human annotator to annotate images presented on the display; An annotation training system, the annotation training system: selecting image samples from the database to be displayed on the display for annotation by the human annotator; receiving one or more suggested annotations from the annotation system; comparing the one or more suggested annotations of the human annotator with previous annotations of the image samples by the expert human annotator or the machine learning annotator; evaluating the annotation performance of the human annotator using a weighting function and a numerical metric based on the results of the comparison; if the annotation performance is below a preset threshold, presenting an attention map on the display to highlight attention areas to draw the attention of the human annotator to any annotation errors identified by the comparison, wherein the annotation errors include providing a personalized explanation of the annotation errors on the display using the attention map; selecting a next image sample from the data block based on any errors identified in the comparison, the next image sample having an attention area of the same type as the attention map, for reinforcement training on the annotation errors identified by the comparison.
8. The system according to claim 7, characterized in that The annotation training system also presents image samples for annotation by the human annotator when the human annotator is evaluated as having annotation performance above a threshold, and adds the annotated image samples from the human annotator to the database.
9. The system according to claim 8, characterized in that The annotated image samples from the human annotators added to the database include weights based on the annotation performance of the human annotators.
10. The system according to claim 7, wherein: The annotation training system also certifies the human annotator for future annotation tasks when the human annotator's annotation performance is above a predetermined level for the type of annotation for which the human annotator has been trained.
11. The system according to claim 7, wherein: The annotation training system also compares annotation performance of multiple human annotators on the same set of images to establish a quality metric for the multiple human annotators.
12. The system according to claim 7, wherein: The image to be annotated includes at least one of a medical image, a geographical image, and an industry image.
13. A computer-readable medium, characterized in that storing computer instructions for training a human annotator to annotate images, which, when executed by one or more processors, cause the one or more processors to: presenting an image sample to a human annotator for annotation, wherein the image sample has been previously annotated by at least one of an expert human annotator and a machine learning annotator; receiving one or more suggested annotations from the human annotator; comparing the one or more suggested annotations of the human annotator with previous annotations of the image sample by the expert human annotator or the machine learning annotator; evaluating the annotation performance of the human annotator using a weighting function and a numerical metric based on a result of the comparison; If the annotation performance is below a preset threshold, presenting an attention map to highlight attention areas to draw the attention of the human annotator to annotation errors identified by the comparison; wherein the annotating the errors comprises providing a personalized explanation of the annotation errors on a display using the attention map; A next image sample is selected based on any errors identified in the comparison, the next image sample having the same type of attention region as the attention map for reinforcement training for annotation errors identified by the comparison.
14. The medium according to claim 13, characterized in that Also included are instructions that, when executed by the one or more processors, cause the one or more processors to present image samples for annotation by the human annotator when the human annotator is evaluated as having annotation performance above a threshold, and to add the annotated image samples from the human annotator to a pool of image samples that includes image samples previously annotated by the expert human annotator or the machine learning annotator.
15. The medium according to claim 14, characterized in that The annotated image samples from the human annotator added to the pool of image samples include weights based on the annotation performance of the human annotator.
Citation Information
Patent Citations
Image report annotation identification
CN106796621A