Consensus Labeling in Digital Pathology Images

By determining consensus locations and labels for annotations within digital pathology images, the technology resolves annotation conflicts, enhancing the accuracy and consistency of ground truth labels, thus improving the performance of machine learning models in digital pathology.

JP2025541735APending Publication Date: 2025-12-23VENTANA MEDICAL SYSTEMS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025531638
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-12-02
Filing Date
2023-11-27
Publication Date
2025-12-23

AI Technical Summary

Technical Problem

Existing techniques for generating consensus labels in digital pathology images fail to address the issue of annotation conflicts, particularly when multiple annotators provide inconsistent or inconsistent annotations, leading to decreased classification accuracy and training machine learning models, resulting in inconsistent performance during inference phase.

Method used

The technology described herein addresses the issue by determining a consensus location and label for a set of annotations within an image, using an annotation processing system that resolves conflicts in annotation locations and labels, thereby improving the accuracy and consistency of ground truth labels for training machine learning models.

Benefits of technology

This approach enhances the accuracy and consistency of ground truth labels, leading to improved performance of machine learning models in digital pathology by mitigating errors in annotation and improving the training process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025541735000001
    Figure 2025541735000001
  • Figure 2025541735000002
    Figure 2025541735000002
  • Figure 2025541735000003
    Figure 2025541735000003
Patent Text Reader

Abstract

The present disclosure relates to systems and methods for determining consensus positions and labels for a set of annotations associated with an object or region within an image. An annotation processing system can access multiple annotations associated with an image depicting at least a portion of a biological specimen. The annotation processing system can determine a consensus position for the set of annotations located at different positions within the region of the image. At the determined consensus position, a consensus label can be determined for the set of annotations that identify target types of different biological structures. The consensus labels across the different positions can be used to generate ground truth labels for the image. The ground truth labels can be used to train a machine learning model configured to predict different types of biological structures within a digital pathology image.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Priority claims This application claims priority to U.S. Provisional Patent Application No. 63 / 385,837, filed December 2, 2022, entitled "CONSENSUS LABELING IN DIGITAL PATHOLOGY IMAGES," which is incorporated by reference in its entirety.

[0002] The present disclosure relates generally to generating ground truth labels for images used in training, validating, or testing machine learning models. More specifically, but without limitation, the present disclosure relates to generating consensus ground truth labels for images. [Background technology]

[0003] In a digital pathology platform, a machine learning model can be used to predict the cell type of various cells depicted in a slide image. Target cell types of cells can include tumor cells, immune cells, stromal cells, etc. The machine learning model can be trained to predict cell types based on a training dataset including a plurality of training images. Each training image can include a set of ground truth labels, which identify the target cell type of cells at the corresponding location in the training image. The set of ground truth labels can include one or more target classifications such that the machine learning model learns to predict one of the classifications of various cellular targets detected in the slide image. The set of ground truth labels is typically assigned by one or more experts (e.g., pathologists).

[0004] However, assigning accurate ground truth labels can be difficult. For example, a first pathologist may consider a particular cell to correspond to a first cell type (e.g., an immune cell), while a second pathologist may consider the same particular cell to correspond to a second cell type (e.g., a tumor cell). These differing observations may spread to all other cells depicted in the training images, thus resulting in training data containing conflicting ground truth labels. Such training data may contain further annotation conflicts because additional experts can be used to assign ground truth labels across various training images. If such training data is used for training, the machine learning model may not learn properly and therefore may not perform consistently during the inference phase (i.e., during deployment). As a result, the classification accuracy of the trained machine learning model may decrease. Summary of the Invention

[0005] The technology described herein relates to determining a consensus location and label for a set of annotations associated with an object or region within an image. An annotation processing system can access multiple annotations associated with an image depicting at least a portion of a biological sample. The annotation processing system can determine a consensus location for the set of annotations located at different locations within the region of the image. At the determined consensus location, a consensus label can be determined for the set of annotations that identify target types of different biological structures. The consensus labels across the different locations can be used to generate a ground truth label for the image.

[0006] In some embodiments, a system is provided that includes one or more data processors and a non-transitory computer-readable storage medium that includes instructions that, when executed on the one or more data processors, cause the one or more data processors to perform some or all of one or more of the methods disclosed herein.

[0007] In some embodiments, a computer program product is provided that is tangibly embodied in a non-transitory machine-readable storage medium and includes instructions configured to cause one or more data processors to perform some or all of one or more of the methods disclosed herein.

[0008] The terms and expressions which have been employed are used as terms of description rather than of limitation, and there is no intention in the use of such terms and expressions to exclude equivalents of the features shown and described or portions thereof, but it is recognized that various modifications are possible within the scope of the invention as claimed. Thus, although the claimed invention has been specifically disclosed by embodiments and optional features, it will be understood that modifications and variations of the concepts disclosed herein may be employed by those skilled in the art, and that such modifications and variations are deemed to be within the scope of the invention as defined by the appended claims. [Brief explanation of the drawings]

[0009] The patent or application file contains at least one drawing executed in color. Copies of this patent or patent application publication with color drawing(s) will be provided by the Office upon request and payment of the necessary fee.

[0010] Aspects and features of various embodiments will become more apparent by way of example only and with reference to the accompanying drawings.

[0011] [Figure 1] FIG. 1 shows an exemplary schematic diagram illustrating a process for generating consensus labels for slide images, according to some embodiments.

[0012] [Figure 2] 1 illustrates an exemplary network for generating digital pathology images, according to some embodiments.

[0013] [Figure 3] 1 illustrates an exemplary computing environment for processing digital pathology images using machine learning / deep learning models, according to various embodiments.

[0014] [Figure 4] FIG. 1 shows a schematic diagram illustrating a process for identifying clusters for a set of annotations of slide images, according to some embodiments.

[0015] [Figure 5] 1 illustrates an exemplary image including annotation locations, according to some embodiments.

[0016] [Figure 6] 1 illustrates a set of exemplary annotations identified in training slide images, according to some embodiments.

[0017] [Figure 7] 10A-10C show exemplary images illustrating the process by which annotation conflicts are resolved, according to some embodiments.

[0018] [Figure 8] 10 shows another set of example images illustrating the process by which annotation conflicts are removed, according to some embodiments.

[0019] [Figure 9] 1 illustrates a process for identifying image annotation-location conflicts, according to some embodiments.

[0020] [Figure 10]1 illustrates a process for removing annotation-location conflicts and determining a consensus label for an image, according to some embodiments.

[0021] [Figure 11] FIG. 1 shows a schematic diagram illustrating a process for removing annotation-label conflicts and determining a consensus label for an image, according to some embodiments.

[0022] [Figure 12] 1 illustrates an example image including a consensus label and location, according to some embodiments.

[0023] [Figure 13] FIG. 1 shows a schematic diagram illustrating a process for removing annotation-label conflicts and determining a consensus label for an image, according to some embodiments. DETAILED DESCRIPTION OF THE INVENTION

[0024] Existing techniques can generally eliminate annotation conflicts by (i) receiving annotations from multiple pathologists for a given image region and (ii) determining a single label for the image region based on the majority of annotations that indicate a particular type of biological structure. However, the above existing techniques can be ineffective when it is unclear whether one or more annotations correspond to the same image region or image object (e.g., a cell). For example, a pathologist may be instructed to annotate a cell by placing a dot at the center of the cell, but annotations may be unintentionally added (i.e., placed) at various locations that do not correspond to the center of the cell. In another example, some annotations may not be added to an image region at all because one or more pathologists may consider the image region not to depict a cell. Therefore, it is difficult to determine an accurate ground truth label for an image based on processing the annotations because annotations are added inconsistently or at different locations on the image.

[0025] Certain embodiments described herein can address these and other issues by determining a consensus location for a set of annotations located at different locations within a region of an image. The images can be part of a training dataset used to train a machine learning model. In some examples, the images are part of a validation dataset used to validate the machine learning model. Additionally or alternatively, the images are part of a test dataset used to test the machine learning model. For the determined consensus location, a consensus label can be determined for the set of annotations that identify different target types and / or locations of biological structures. The consensus labels across the different locations can be used to generate ground truth labels for the image. The ground truth labels can be used to train a machine learning model configured to predict different types and / or locations of biological structures within a digital pathology image. In effect, determining consensus locations and labels can increase the consistency and accuracy of generating ground truth labels for training machine learning models.

[0026] The annotation processing system can access multiple annotations associated with an image depicting at least a portion of a biological sample. As an illustrative example, the image can be a digital pathology image depicting at least a portion of a lung tissue sample obtained from a subject. The multiple annotations can be used to generate ground truth labels for supervised learning of a machine learning model, where each annotation can include (i) a location identifying an object or region within the image, (ii) a target type of biological structure for the object or region, and (iii) an identifier associated with the annotator (e.g., a pathologist) who generated the annotation. Continuing with the above example, the annotation processing system can receive annotations generated for an image by a team of pathologists (e.g., four pathologists). The annotations in this example can include a set of x-y coordinates identifying the location of the image region, stained tumor cells (e.g., TC+) as the target type of biological structure, and "Pathway 4" as the identifier of the annotator who generated the annotation.

[0027] The annotation processing system may determine annotation-location conflicts for one or more annotations associated with a region or object of an image. In some examples, annotation-location conflicts are determined based on a region having two or more annotations from a single annotator. Continuing the example above, the annotation processing system may identify an annotation-location conflict for a particular region by determining that five annotations are associated with the particular image region and further determining that two of the five annotations were submitted by an annotator identified as "Path 4."

[0028] The annotation processing system can then resolve annotation location conflicts for regions or objects of the image. For each of two or more annotations of a region associated with the same annotator, the annotation processing system can determine a statistical value (e.g., mean, median) representing the distance from the annotation to other annotations in the corresponding cluster. The statistical values ​​between two or more annotations from the same annotator can be compared. The annotation with the lowest statistical value can be determined to be part of the region, while the remaining annotations in the conflict can be excluded from being part of the region. Additionally or alternatively, the annotation processing system can apply a clustering algorithm (e.g., a k-means algorithm) to group the annotations so that a first subset of annotations represents a sub-region of the first region and a second subset of annotations represents a sub-region of the second region of the image.

[0029] Continuing with the above example, the annotation processing system may determine a first statistical value of 10 pixels for a first annotation generated by "Path 4," where the first statistical value corresponds to the average distance between the first annotation and other annotations present in a particular region. The annotation processing system may also determine a second statistical value of 15 pixels for a second annotation generated by "Path 4," where the second statistical value corresponds to the average distance between the second annotation and other annotations. The annotation processing system may resolve annotation location conflicts by discarding the second annotation with the larger statistical value.

[0030] After a consensus position is determined for the set of annotations, the annotation processing system can determine a ground truth label for the consensus position of the image. In some examples, a ground truth label is determined when all of the set of annotations indicate the same biological structure target type (e.g., tumor cell). However, there are instances where there is no clear consensus among the sets of annotations regarding the biological structure target type. In such cases, the annotation processing system can resolve the "annotation-label" conflict by (i) determining a weight associated with each annotation and (ii) comparing the weights to determine a consensus label for the consensus position. The weight is determined based on performance and historical data associated with the annotator who generated the annotation. For example, the annotator of a first annotation may be a senior pathologist with many years of experience. The weight of the first annotation may be assigned a relatively higher value than the weight of another, second annotation associated with a junior pathologist.

[0031] Continuing with the example, a consensus location includes four annotations: annotators "Path 1" and "Path 3" label a location as having stained tumor cells, while two other annotators, "Path 2" and "Path 4," label the location as having unstained tumor cells. The sum of the first weights assigned to "Path 1" and "Path 3" can then be compared with the sum of the second weights assigned to "Path 2" and "Path 4." If the first sum is determined to be greater than the second sum, the annotation processing system can determine a consensus label for the consensus location as "stained tumor cells."

[0032] The annotation processing system can then generate a set of ground truth labels for the image based at least in part on the consensus positions and corresponding consensus labels. The set of ground truth labels for the image can then be used to perform supervised learning of the machine learning model. In some examples, the set of ground truth labels for the image is used to validate the machine learning model. Additionally or alternatively, the set of ground truth labels for the image can be used to test the machine learning model. By determining the consensus positions and labels, accurate and consistent performance of the machine learning model can be achieved.

[0033] In some examples, the annotation processing system optimizes the determination of the consensus label based on various adjustable parameters. For example, the annotation processing system may allow adjustment of weights based on each annotator's performance across different images or the FOV of a single image. In addition to adjusting weights, other types of customizations may be applied by a user to determine the consensus location and label. For example, a user may adjust a distance threshold to determine whether a set of annotations identifies similar locations on an image. These customizations allow for a better representation of the data on which the consensus is calculated. Additionally, annotations and any conflicts between annotations may be displayed in a graphical user interface, allowing a user to adjust weight values ​​and distance thresholds.

[0034] Certain embodiments described herein improve the training of machine learning models that classify biological structures in digital pathology images. An annotation processing system can improve the consistency and accuracy of ground truth label generation by facilitating the resolution of various types of conflicts between annotations across images (e.g., annotation-location conflicts, annotation-label conflicts). In effect, the annotation processing system can mitigate errors made by annotators (e.g., pathologists), such as clicking a particular annotation in the wrong location on an image when generating an annotation. Furthermore, the use of adjustable weights for annotators can fine-tune the consistency and accuracy of how ground truth labels for images are generated. Thus, embodiments herein reflect improved capabilities of artificial intelligence systems and digital pathology image processing technologies.

[0035] While specific embodiments are described, these embodiments are presented by way of example only and are not intended to limit the scope of protection. The apparatus, methods, and systems described herein may be embodied in various other forms. Furthermore, various omissions, substitutions, and changes in form of the exemplary methods and systems described herein may be made without departing from the scope of protection.

[0036] I. Definition As used herein, when an action is "based on" something, this means that the action is based at least in part on at least a part of the something.

[0037] As used herein, the terms "substantially," "approximately," and "about" are defined as generally, but not necessarily exactly, what is specified (including exactly what is specified), as would be understood by one of ordinary skill in the art. In any disclosed embodiment, the terms "substantially," "approximately," or "about" may be replaced with "within [percentage]" of what is specified, where percentages include 0.1, 1, 5, and 10%.

[0038] As used herein, the terms "sample," "biological sample," "tissue," or "tissue sample" refer to any sample containing biomolecules (such as proteins, peptides, nucleic acids, lipids, carbohydrates, or combinations thereof) obtained from any organism, including viruses. Other examples of organisms include mammals (e.g., veterinary animals such as humans, cats, dogs, horses, cows, and pigs, and laboratory animals such as mice, rats, and primates), insects, annelids, arachnids, marsupials, reptiles, amphibians, bacteria, and fungi. Biological samples include tissue samples (such as tissue sections and needle biopsies of tissue), cell samples (such as cytological smears, such as Pap smears or blood smears, or samples of cells obtained by microdissection), or cell fractions, fragments, or organelles (such as those obtained by lysing cells and separating their components, such as by centrifugation). Other examples of biological samples include blood, serum, urine, semen, feces, cerebrospinal fluid, interstitial fluid, mucus, tears, sweat, pus, biopsy tissue (e.g., obtained by surgical or needle biopsy), nipple aspirate, earwax, milk, vaginal fluid, saliva, swabs (such as oral swabs), or any material containing biomolecules derived from an initial biological sample. In some embodiments, the term "biological sample," as used herein, refers to a sample (such as a homogenized or liquefied sample) prepared from a tumor or portion thereof obtained from a subject.

[0039] As used herein, the terms "biological material," "biological structure," or "cellular structure" refer to naturally occurring materials or structures that comprise all or part of a biological structure (e.g., a cell nucleus, cell membrane, cytoplasm, chromosomes, DNA, cell, cell mass, etc.).

[0040] As used herein, "digital pathology image" refers to a digital image of a stained sample.

[0041] As used herein, the term "cell detection" refers to the detection of the location and characteristics of a cell or cell structure (e.g., cell nucleus, cell membrane, cytoplasm, chromosome, DNA, cell, cell mass, etc.) pixel.

[0042] As used herein, the term "region" or "object" refers to a group of pixels in an image that contains image data intended to be evaluated in an image analysis process. A region or object may depict a tissue area intended to be analyzed in the image analysis process (e.g., tumor cells or staining representations).

[0043] As used herein, the term "field of view" or "FOV" refers to an image of an entire slide or a portion of an entire slide. In some embodiments, "FOV" refers to an area of ​​the entire slide scan or a region of interest having (x,y) pixel dimensions (e.g., 1000 pixels by 1000 pixels). In some embodiments, the FOV includes the same pixel resolution and / or magnification level associated with a corresponding entire slide image. Additionally or alternatively, the FOV includes different pixel resolutions and / or magnification levels associated with a corresponding entire slide image.

[0044] II. Overview 1 shows an exemplary schematic 100 illustrating a process for generating consensus labels for slide images, according to some embodiments. The schematic 100 includes (i) an annotation processing system 102 configured to generate a consensus location and label for a given image, and (ii) a graphical user interface 116 configured to display intermediate output of the annotation processing system 102 to a user.

[0045] In block 104, the annotation processing system 102 receives a set of annotations associated with an image. The image may be a digital pathology image depicting at least a portion of a biological sample (e.g., tissue) obtained from a subject. In some examples, an annotation tool is used by an annotator to provide a set of annotations for the entire tissue image or a selected region of the image (e.g., field of view (FOV)). The annotation tool may aggregate the set of annotations into a schema file (e.g., an XML document) and send the schema file to the annotation processing system 102 for further processing. In some examples, an annotation tool is provided to each of multiple annotators involved in generating ground truth labels for the image, and each schema file is sent to the annotation processing system 102 for further processing. For example, the annotation processing system 102 may receive five schema files, each containing a set of annotations submitted by a corresponding annotator (e.g., a pathologist). In some embodiments, the annotation processing system 102 receives and processes 10, 15, 20, 25, 30, 35, 40, 45, 50, or more than 50 schema files for the corresponding image.

[0046] The annotation tool can also provide labels that can be used to classify cells in a given image. Annotations can be added to locations in the image, and the annotations can include labels that identify the type of cell depicted in the image. In some examples, the labels include (i) a TC+ label that identifies tumor cells stained with a specific biomarker; (ii) a TC label that identifies other unstained tumor cells; (iii) an IC label that identifies immune cells; and (iv) other labels that identify cells other than tumor cells or immune cells.

[0047] In response to receiving a set of annotations, the annotation processing system 102 may analyze each annotation in the set to identify its corresponding annotation data. For example, the annotations may be analyzed by the annotation processing system 102 to identify an initial location of a region of the image and an initial label (e.g., a stained tumor cell) that classifies a biological structure depicted in the initial location of the image. In addition to the initial location and label, the annotation processing system 102 may identify an identifier of the annotator who submitted the annotation, the type of tissue depicted by the image, the number of annotators who submitted a set of annotations for the image, and the label used to classify the cells in the image.

[0048] In block 106, the annotation processing system 102 identifies one or more clusters of annotations associated with the image. For example, the location of the annotations may include a set of coordinates of the image that the annotator indicated to depict a particular biological structure. A cluster may be defined as a set of annotations that are within a predetermined distance. In some examples, a threshold is used to determine the set of annotations, the threshold identifying a number of pixels (e.g., 10 pixels) between the annotations. The threshold may be adjusted by the annotator, which may increase or decrease the number of annotations included in the cluster. Adjusting the threshold to increase or decrease the number of annotations may define the size of the cluster for the image.

[0049] In block 108, the annotation processing system 102 identifies annotation location conflicts for the annotations. The annotation processing system 102 may identify annotation location conflicts by determining that two or more annotations associated with the same annotator (e.g., a pathologist) are associated within a particular cluster. Additionally or alternatively, the annotation processing system 102 may identify annotation location conflicts by detecting that the number of annotations in a particular cluster (e.g., six annotations) exceeds the total number of annotators who submitted annotations for the image (e.g., five pathologists). In some examples, the graphical user interface 116 displays annotations with annotation-location conflicts (block 118).

[0050] In block 110, the annotation processing system 102 resolves annotation location conflicts. For annotations that include location conflicts, the annotation processing system 102 may determine, for each of two or more annotations with conflicts (e.g., annotations generated by a particular annotator for the same object), a statistical value (e.g., mean, median) that represents the distance from the annotation to other annotations within the corresponding region. The statistical values ​​between two or more annotations from the same annotator may be compared. The annotation with the lowest statistical value may be determined to be part of the region's annotations, and the remaining conflicting annotations may be exempt from being associated with the region. In some examples, the graphical user interface 116 displays the annotations with resolved annotation-location conflicts (block 120).

[0051] In block 112, the annotation processing system 102 determines a consensus location for one or more regions of the image. Each region may include a set of annotations submitted by multiple annotators, and the consensus location for the cluster may be determined based on the median coordinate between the locations of the sets of annotations associated with the cluster. In some examples, the count of annotations for the region is compared to a predetermined confidence threshold. If the count falls below the confidence threshold, the set of annotations is discarded from the image. As a result, no ground truth label is generated for such region.

[0052] In block 114, the annotation processing system 102 determines a consensus label for the image. The annotation processing system 102 can use value weighting to determine a consensus label even if a majority of labels are not present for a given consensus location. For example, a consensus location may include four annotations, two of which include labels identifying locations depicting tumor cells and two of which include other labels identifying locations depicting normal cells. Based on the weights applied to each of the annotations, the annotation processing system 102 can resolve "ties" between the annotations and determine a consensus label for the location. In some examples, the graphical user interface 116 displays the consensus location and label for the image (block 122).

[0053] In some examples, the graphical user interface 116 allows adjustment of one or more parameters used by the annotation processing system 102 in response to one or more of the visualizations 118, 120, and 122. For example, the graphical user interface 116 may adjust a weighted value assigned to an annotator associated with a particular set of annotations. The weighted value may be adjusted by determining an accuracy metric for the annotator based on verification of images with ground truth labels. The accuracy metric may be determined based on the percentage of annotations submitted by the annotator that correspond to the ground truth (e.g., 75%, 85%). In some examples, the annotator's weighted value is adjusted to be directly proportional to the accuracy metric determined for the annotator.

[0054] III. Generation of digital pathology images Digital pathology involves the interpretation of digitized images to accurately diagnose a subject and guide treatment decisions. In a digital pathology solution, an image analysis workflow can be established to automatically detect or classify biological objects of interest, such as positive or negative tumor cells. An exemplary digital pathology solution workflow includes acquiring a tissue slide, scanning a preselected region or the entire tissue slide with a digital image scanner (e.g., a whole slide imaging (WSI) scanner) to obtain a digital image, performing image analysis on the digital image using one or more image analysis algorithms, and potentially detecting and quantifying each object of interest (e.g., counting or identifying its specific or cumulative area) based on the image analysis (e.g., quantitative or semi-quantitative scoring such as positive, negative, moderate, or weak).

[0055] 2 shows an exemplary network 200 for generating digital pathology images. A fixation / embedding system 205 fixes and / or embeds tissue samples (e.g., samples containing at least a portion of at least one tumor) using a fixative (e.g., a liquid fixative such as a formaldehyde solution) and / or an embedding substance (e.g., a histology wax such as paraffin wax, and / or one or more resins such as styrene or polyethylene). Each sample may be fixed by exposing the sample to a fixative for a predetermined period of time (e.g., at least 3 hours) and then dehydrating the sample (e.g., via exposure to an ethanol solution and / or a clearing intermediate agent). The embedding substance may become infiltrated when the sample is in a liquid state (e.g., upon heating).

[0056] Fixation and / or embedding of samples is used to preserve samples and slow their degradation. In histology, fixation generally refers to an irreversible process using chemicals to preserve chemical composition, preserve natural sample structure, and protect cellular structures from degradation. Fixation may also harden cells or tissues for sectioning. Fixatives may enhance sample and cell preservation by cross-linking proteins. Fixatives may bind to and cross-link some proteins and denature others through dehydration, which hardens tissue and can inactivate enzymes that might otherwise degrade the sample. Fixatives may also kill bacteria.

[0057] Fixatives can be administered, for example, by perfusion and immersion of the prepared sample. Various fixatives can be used, including methanol, buprenorphine fixatives, and / or formaldehyde fixatives, such as neutral buffered formalin (NBF) or paraffin-formalin (paraformaldehyde-PFA). If the sample is a liquid (e.g., blood sample), the sample may be smeared onto a slide and allowed to dry before fixation. While the fixation process can help preserve the structure of the sample and cells for histological examination, fixation can mask tissue antigens, thereby reducing antigen detection. Therefore, fixation is generally considered a limiting factor in immunohistochemistry, as formalin can crosslink antigens and mask epitopes. In some instances, additional processes are performed to reverse the effects of crosslinking, including treating the fixed sample with citraconic anhydride (a reversible protein crosslinker) and heating.

[0058] Embedding can involve infiltrating a sample (e.g., a fixed tissue sample) with an appropriate histological wax, such as paraffin wax. Histological waxes can be insoluble in water or alcohol but soluble in paraffin solvents such as xylene. Therefore, it may be necessary to replace the water in the tissue with xylene. To do so, the sample can first be dehydrated by gradually replacing the water in the sample with alcohol. This can be achieved by passing the tissue through increasing concentrations of ethyl alcohol (e.g., from 0 to approximately 100%). After replacing the water with alcohol, the alcohol can be replaced with xylene, which is miscible with alcohol. Because histological waxes can be soluble in xylene, the molten wax can be filled with xylene, filling the spaces previously filled with water. The wax-filled sample can be cooled to form a hardened block, which can be clamped to a microtome, vibratome, or compressome for sectioning. In some cases, deviations from the exemplary procedure above can result in infiltration of the paraffin wax, inhibiting penetration of antibodies, chemicals, or other fixatives.

[0059] The tissue slicer 210 can then be used to section the fixed and / or embedded tissue sample (e.g., a tumor sample). Sectioning is the process of cutting thin slices (e.g., 2-5 μm thick) of a sample from a tissue block for mounting on a microscope slide for examination. Sectioning may be performed using a microtome, vibratome, or compressome. In some cases, tissue can be rapidly frozen in dry ice or isopentane and then cut with a cold knife in a refrigerated cabinet (e.g., a cryostat). Other types of coolants, such as liquid nitrogen, can be used to freeze tissue. Sections for use with brightfield and fluorescence microscopy are generally on the order of 2-10 μm thick. In some cases, sections can be embedded in epoxy or acrylic resin, which may allow for cutting thinner sections (e.g., <2 μm). These sections may then be mounted on one or more glass slides. A coverslip may be placed on top to protect the sample sections.

[0060] Because tissue sections and the cells therein are substantially transparent, slide preparation typically further includes staining the tissue sections (e.g., automated staining) to make relevant structures more visible. In some cases, the staining is performed manually. In other cases, the staining is performed semi-automatically or automatically using a staining system 215. The staining process involves exposing sections of the tissue sample or fixed liquid sample to one or more different stains (e.g., sequentially or simultaneously) to reveal different properties of the tissue.

[0061] For example, stains can be used to mark specific types of cells and / or flag specific types of nucleic acids and / or proteins to aid in microscopy. The staining process generally involves adding a dye or stain to a sample to confirm or quantify the presence of a particular compound, structure, molecule, or feature (e.g., subcellular feature). For example, stains can help identify or highlight specific biomarkers in tissue sections. In other examples, stains can be used to identify or highlight biological tissues (e.g., muscle fibers or connective tissue), cell populations (e.g., different blood cells), or organelles within individual cells.

[0062] One exemplary type of tissue stain is a histochemical stain, which uses one or more chemical dyes (e.g., acid dyes, basic dyes, chromogens) to stain tissue structures. Histochemical stains can be used to reveal general aspects of tissue morphology and / or cellular microanatomy (e.g., distinguishing cell nuclei from cytoplasm, revealing lipid droplets, etc.). One example of a histochemical stain is H&E. Other examples of histochemical stains include trichrome stains (e.g., Masson's trichrome), periodic acid-Schiff (PAS), silver stains, and iron stains. The molecular weight of histochemical stains (e.g., dyes) is generally about 500 kilodaltons (kD) or less, although some histochemical stains (e.g., Alcian blue, phosphomolybdic acid (PMA)) can have molecular weights of up to 2000 or 3000 kD. One example of a high-molecular-weight histochemical stain is α-amylase (about 55 kD), which may be used to reveal glycogen.

[0063] Another type of tissue staining is IHC, also known as "immunostaining," which uses a primary antibody that specifically binds to a target antigen of interest (also known as a biomarker). IHC can be direct or indirect. In direct IHC, the primary antibody is directly conjugated to a label (e.g., a chromophore or fluorophore). In indirect IHC, a primary antibody is first bound to the target antigen, and then a secondary antibody conjugated to a label (e.g., a chromophore or fluorophore) is bound to the primary antibody. Because antibodies have a molecular weight of approximately 150 kD or greater, the molecular weight of IHC reagents is much larger than that of histochemical staining reagents.

[0064] Various types of staining protocols may be used to perform staining. For example, an exemplary IHC staining protocol includes using a hydrophobic barrier around the sample (e.g., tissue section) to prevent leakage of reagents from the slide during incubation, treating the tissue section with reagents to block endogenous sources of nonspecific staining (e.g., enzymes, free aldehyde groups, immunoglobulins, and other unrelated molecules that may mimic specific staining), incubating the sample with a permeabilization buffer to promote penetration of antibodies and other staining reagents into the tissue, incubating the tissue section with a primary antibody at a specific temperature (e.g., room temperature, 6-8 °C) for a certain period of time (e.g., 1-24 hours), rinsing the sample using a wash buffer, and then incubating the sample (tissue section) with a secondary antibody at another specific temperature (e.g., room temperature) for another period of time, rinsing the sample again using a water buffer, incubating the rinsed sample with a chromogen (e.g., DAB: 3,3'-diaminobenzidine), and washing away the chromogen to stop the reaction. In some instances, a counterstain is then used to distinguish the overall "landscape" of the sample and serve as the primary color reference used to detect tissue targets. Counterstains may include, for example, hematoxylin (a blue-to-purple stain), methylene blue (a blue stain), toluidine blue (a stain that renders nuclei deep blue and polysaccharides pink to red), nuclear fast red (also known as Kern Echtrot dye, a red stain), methyl green (a green stain), and non-nucleogenic stains such as eosin (a pink stain). Those skilled in the art will recognize that staining can be performed using other immunohistochemical staining techniques.

[0065] In another example, an H&E staining protocol can be performed to stain tissue sections. The H&E staining protocol involves applying hematoxylin stain mixed with a metal salt or mordant to the sample. The sample can then be rinsed in a weak acid solution to remove excess stain (differentiation), followed by bluing in weak alkaline water. After application of hematoxylin, the sample can be counterstained with eosin. It will be understood that other H&E staining techniques can be performed.

[0066] In some embodiments, staining can be performed using various types of staining agents depending on the feature of interest. For example, DAB can be used on various tissue sections for IHC staining, producing a brown color that depicts the feature of interest in the stained image. In another example, alkaline phosphatase (AP) can be used on skin tissue sections for IHC staining because the DAB color can be masked by melanin pigments. Regarding primary staining techniques, applicable staining agents can include, for example, basophilic and eosinophilic stains, hematin and hematoxylin, silver nitrate, trichrome stains, etc. Acidic dyes can react with cationic or basic components in tissues or cells, such as proteins and other components in the cytoplasm. Basic dyes can react with anionic or acidic components in tissues or cells, such as nucleic acids. As mentioned above, one example of a staining system is H&E. Eosin can be a negatively charged pink acidic dye, and hematoxylin can be a purple or blue basic dye containing hematin and aluminum ions. Other examples of stains may include Periodic Acid-Schiff (PAS) stain, Masson's trichrome, Alcian blue, Van Gieson, reticulin stain, etc. In some embodiments, different types of stains may be used in combination.

[0067] The sections may then be mounted on corresponding slides, and the imaging system 220 may then scan or image them to generate raw digital pathology images 225a-n. A microscope (e.g., an electron microscope or an optical microscope) may be used to magnify the stained specimen. For example, an optical microscope may have a resolution of less than 1 μm, such as on the order of several hundred nanometers. To observe finer details in the nanometer or sub-nanometer range, an electron microscope may be used. An imaging device (combined with or separate from the microscope) images the magnified biological specimen to acquire image data, such as a multichannel image (e.g., multichannel fluorescence) having several (e.g., 10-16) channels. The imaging device may include, but is not limited to, a camera (e.g., an analog camera, a digital camera, etc.), optical elements (e.g., one or more lenses, a sensor-focusing lens group, a microscope objective, etc.), an imaging sensor (e.g., a charge-coupled device (CCD), a complementary metal-oxide semiconductor (CMOS) image sensor, etc.), photographic film, etc. In digital embodiments, the imaging device may include multiple lenses that cooperate to achieve on-the-fly focusing. An image sensor, such as a CCD sensor, may capture a digital image of the biological sample. In some embodiments, the imaging device is a bright-field imaging system, a multispectral imaging (MSI) system, or a fluorescence microscope system. The imaging device may utilize invisible electromagnetic radiation (e.g., UV light) or other imaging techniques to capture the image. For example, the imaging device may include a microscope and a camera configured to capture an image magnified by the microscope. The image data received by the analysis system may be identical to and / or derived from the raw image data captured by the imaging device.

[0068] Images of the stained sections may then be stored on a storage device 225, such as a server. The images may be stored locally, remotely, and / or on a cloud server. Each image may be stored in association with a subject identifier and a date (e.g., the date the sample was collected and / or the date the image was captured). The images may further be transmitted to another system (e.g., a system associated with a pathologist, an automated or semi-automated image analysis system, or a machine learning training and deployment system, as described in further detail herein).

[0069] It will be understood that modifications to the process described with respect to network 200 are contemplated. For example, if the sample is a liquid sample, embedding and / or sectioning may be omitted from the process.

[0070] IV. Exemplary Systems for Digital Pathology Image Conversion 3 shows a block diagram illustrating a computing environment 300 for processing digital pathology images using a machine learning model. As described further herein, processing the digital pathology image may include training a machine learning algorithm using the digital pathology image and / or converting some or all of the digital pathology image into one or more results using a trained (or partially trained) version of the machine learning algorithm (i.e., a machine learning model).

[0071] As shown in FIG. 3, the computing environment 300 includes several stages: an image storage stage 305, a pre-processing stage 310, a labeling stage 315, a data augmentation stage 317, a training stage 320, and a result generation stage 325.

[0072] The image storage stage 305 includes one or more image data stores 330 (e.g., storage device 230 described in connection with FIG. 2 ) that are accessed (e.g., by the pre-processing stage 310) to provide a set of digital images 335 of pre-selected areas from a biological specimen slide (e.g., a histology slide) or the entire biological specimen slide. Each digital image 335 stored in each image data store 330 and accessed by the image storage stage 310 may include a digital pathology image generated according to some or all of the processes described with respect to the network 200 depicted in FIG. 2 . In some embodiments, each digital image 335 includes image data from one or more scanned slides. Each of the digital images 335 may correspond to image data from a single specimen and / or from a single day on which the underlying image data corresponding to the image was collected.

[0073] The image data may include the image, as well as any information about the color or wavelength channels, and details about the imaging platform on which the image was generated. For example, a tissue section may need to be stained by applying a staining assay containing one or more different biomarkers associated with a chromogenic stain for brightfield imaging or a fluorophore for fluorescent imaging. The staining assay may use a chromogenic stain for brightfield imaging, an organic fluorophore for fluorescent imaging, quantum dots, or a combination of organic fluorophores and quantum dots, or any other combination of stains, biomarkers, and observation or imaging devices. Exemplary biomarkers include estrogen receptor (ER), human epidermal growth factor receptor 2 (HER2), human Ki-67 protein, progesterone receptor (PR), programmed cell death protein 1 (PD1), etc., and the tissue section is detectably labeled with a binding agent (e.g., an antibody) for each of the ER, HER2, Ki-67, PR, PD1, etc. In some embodiments, digital image and data analysis operations such as classification, scoring, COX modeling, and risk stratification depend on the type of biomarker used and the selection and annotation of the field of view (FOV). Furthermore, typical tissue sections are processed in an automated staining / assay platform that applies a staining assay to the tissue section, thereby obtaining a stained sample. Various commercial products suitable for use as staining / assay platforms exist on the market, one example being the VENTANA® SYMPHONY® product from the assignee, Ventana Medical Systems, Inc. The stained tissue sections may be fed into an imaging system, such as a microscope or whole-slide scanner with a microscope and / or imaging component, one example being the VENTANA® iScan Coreo® / VENTANA® DP200 product from the assignee, Ventana Medical Systems, Inc. Multiple tissue slides can be scanned with equivalent multiple-slide scanner systems.Additional information provided by the imaging system may include any information regarding the staining platform, including the concentration of chemicals used for staining, the reaction time of the chemicals applied to the tissue in the staining, and / or the pre-analysis conditions of the tissue, such as the age of the tissue, fixation method, duration, embedding method for sections, and cutting method.

[0074] In the preprocessing stage 310, one, more, or all of the set of digital images 335 are preprocessed using one or more techniques to generate corresponding preprocessed images 340. Preprocessing may include cropping the image. In some examples, preprocessing may further include standardizing or rescaling (e.g., normalizing) all features to the same scale (e.g., the same size scale or the same color or saturation scale). In particular cases, the image is resized to a minimum size (width or height) of a predetermined number of pixels (e.g., 2500 pixels) or a maximum size (width or height) of a predetermined number of pixels (e.g., 3000 pixels), optionally maintaining the original aspect ratio. Preprocessing may further include removing noise. For example, the image may be smoothed, such as by applying a Gaussian function or Gaussian blur, to remove unwanted noise.

[0075] The preprocessed images 340 may include one or more training images, validation images, test images, and unlabeled images. It should be understood that the preprocessed images 340 corresponding to the training, validation, and unlabeled groups do not need to be accessed simultaneously. For example, an initial set of training and validation preprocessed images 340 may be initially accessed and used to train the machine learning algorithm 355, and subsequently, unlabeled input images may be accessed or received (e.g., one or more times thereafter) and used by the trained machine learning model 360 to produce a desired output (e.g., cell classification).

[0076] In some examples, the machine learning algorithm 355 is trained using supervised learning, and some or all of the preprocessed image 340 is partially or fully labeled in the labeling stage 315, either manually, semi-automatically, or automatically, with labels 345 that identify the “correct” interpretation (i.e., “ground truth”) of various biological materials and structures within the preprocessed image 340. For example, the labels 345 may identify features of interest (for example), a cellular classification, a binary description of whether a given cell is a particular type of cell, a binary description of whether the preprocessed image 340 (or a particular region of the preprocessed image 340) contains a particular type of description (e.g., necrosis or artifact), slide-level or region-specific categorical features of description (e.g., identifying a particular type of cell), counts (e.g., identifying the amount of a particular type of cell within a region, the amount of depicted artifact, or the amount of necrotic area), the presence or absence of one or more biomarkers, etc. In some cases, the labels 345 include location. For example, the label 345 may identify the point location of the nucleus of a particular type of cell, or the point location of a particular type of cell (e.g., a raw dot label). As another example, the label 345 may include a border of a delineated tumor boundary, blood vessel, necrotic region, etc. As another example, the label 345 may include one or more biomarkers identified based on biomarker patterns observed using one or more stains. For example, a tissue slide stained for a biomarker, such as programmed cell death protein 1 ("PD1"), may be observed and / or processed to label cells as either positive or negative cells, taking into account the expression level and pattern of PD1 in the tissue. Depending on the feature of interest, a given labeled pre-processed image 340 may be associated with a single label 345 or multiple labels 345. In the latter case, each label 345 may be associated with (for example) an indication as to which location or portion within the pre-processed image 345 the label corresponds to.

[0077] The labels 345 assigned in the labeling stage 315 may be identified based on input from a human user (e.g., a pathologist or image scientist) and / or an algorithm (e.g., an annotation tool) configured to define the labels 345. In some examples, the labeling stage 315 may include transmitting and / or presenting some or all of the one or more preprocessed images 340 to a computing device operated by a user. In some examples, the labeling stage 315 includes utilizing an interface presented by a labeling controller 350 (e.g., using an API) at the computing device operated by the user, the interface including an input component for accepting input identifying labels 345 for features of interest. For example, a user interface may be presented by the labeling controller 350 that enables selection of an image or region of an image (e.g., FOV) for labeling. A user operating a terminal may select an image or FOV using the user interface. Several image or FOV selection mechanisms may be provided, such as specifying a known or irregular shape or defining an anatomical region of interest (e.g., tumor region). In one example, the image or FOV is a selected entire tumor region on an IHC slide stained with a combination of H&E stains. The selection of the image or FOV can be performed by a user or by an automated image analysis algorithm, such as tumor region segmentation on an H&E tissue slide. For example, a user can select that an image or FOV of the entire slide or tumor, or the entire slide or tumor region, can be automatically designated as the image or FOV using a segmentation algorithm. The user operating the terminal can then select one or more labels 345 to be applied to the selected image or FOV, such as point locations on cells, positive markers for biomarkers expressed by cells, negative biomarkers for biomarkers not expressed by cells, or a boundary around the cells.

[0078] In some examples, the interface may identify which particular labels 345 are requested and / or to what extent, which may be communicated to the user via (for example) text instructions and / or visualization. For example, a particular color, size, and / or symbol may indicate that a label 345 is requested for a particular representation within an image (e.g., a particular cell or region or staining pattern) relative to other representations. If labels 345 corresponding to multiple representations are requested, the interface may identify each of the representations simultaneously or sequentially (such that labeling one identified representation triggers the identification of the next representation for labeling). In some examples, each image is presented until the user identifies a particular number of labels 345 (e.g., of a particular type). For example, a given whole slide image or a given patch of a whole slide image may be presented until the user identifies the presence or absence of three different biomarkers, at which point the interface may present a different whole slide image or image of a different patch (e.g., until a threshold number of images or patches have been labeled). Thus, in some examples, the interface is configured to request and / or accept labels 345 of an incomplete subset of features of interest, and the user may determine which of the potentially many depictions will be labeled.

[0079] In some examples, the labeling stage 315 includes a labeling controller 350 that implements an annotation algorithm to semi-automatically or automatically label various features of an image or a region of interest within the image. The labeling controller 350 annotates the image or FOV of a first slide according to user input or an annotation algorithm and maps the annotations across the remaining slides. Depending on the defined FOV, several methods for annotation and alignment are possible. For example, the entire tumor region annotated on an H&E slide from multiple consecutive slides can be selected automatically or by the user on an interface such as VIRTUOSO / VERSO™. Because the other tissue slides correspond to serial sections from the same tissue block, the labeling controller 350 performs a marker-to-marker alignment operation to map and transfer the entire tumor annotation from the H&E slide to each of the remaining IHC slides in the series. Exemplary methods for marker-to-marker registration are described in further detail in commonly assigned International Publication No. WO 2014140070, "Whole slide image registration and cross-image annotation devices, systems and methods," filed March 12, 2014, which is incorporated herein by reference in its entirety for all purposes. In some embodiments, any other method for image registration and generation of a whole tumor annotation may be used. For example, a qualified radiologist, such as a pathologist, may annotate the whole tumor region on any other IHC slide and execute the labeling controller 350 to map the whole tumor annotation on other digitized slides. For example, a pathologist (or an automated detection algorithm) may annotate the whole tumor region on an H&E slide and trigger analysis of all adjacent serially sectioned IHC slides to determine a whole-slide tumor score for the annotated regions on all slides.

[0080] In some cases, the labeling stage 315 further includes an annotation processing system 351 that implements an annotation algorithm to identify annotation position and annotation label conflicts within a set of annotations associated with the image (or FOV of the image). The annotation processing system 351 can determine a consensus position for a set of annotations located at different positions within a region of the training image. In some cases, the annotation processing system 351 determines that an annotation position conflict exists for a region in the training image by determining that two or more annotations from the same annotator are present within the region. The annotation processing system 351 can resolve such position conflicts by retaining the annotation that has the closest distance to other annotations in the region while discarding other annotations from the same annotator. With the determined consensus position, a consensus label can be determined for the set of annotations that identify target types of different biological structures. The consensus labels across the different positions can be used to generate ground truth labels for the image. The ground truth labels can be used to train, validate, and / or test machine learning models configured to predict different types of biological structures within digital pathology images.

[0081] In the augmentation stage 317, a training set of labeled or unlabeled images (original images) from the preprocessed images 340 is augmented with synthetic images 352 generated using an augmentation control 354 that executes one or more augmentation algorithms. Augmentation techniques are used to artificially increase the amount and / or type of training data by adding slightly modified synthetic copies of existing training data or newly created synthetic data from existing training data. As described herein, variations between scanners and laboratories can cause intensity and color variability in digital images. Additionally, poor scanning can result in gradient variations and blurring effects, assay staining can result in staining artifacts such as background washout, and different tissue / patient samples can result in cell size variations. These variations and perturbations can adversely affect the quality and reliability of deep learning and artificial intelligence networks. The augmentation techniques implemented in the augmentation stage 317 act as regularizers of these variations and perturbations, helping to reduce overfitting when training machine learning models. It should be understood that the enhancement techniques described herein can be used as regularizers for any number and type of fluctuations and disturbances and are not limited to the various specific examples described herein.

[0082] During the training phase 320, the labels 345 and corresponding preprocessed images 340 can be used by a training controller 365 to train a machine learning algorithm 355 according to various workflows described herein. For example, to train the algorithm 355, the preprocessed images 340 may be partitioned into a subset of images 340a (e.g., 90%) for training and a subset of images 340b (e.g., 10%) for validation. The partitioning may be performed randomly (e.g., 90 / 10% or 70 / 30%), or according to more complex validation techniques such as K-fold cross-validation, leave-one-out cross-validation, leave-one-out cross-validation, nested cross-validation, etc., to minimize sampling bias and overfitting. The partitioning may also be performed based on the inclusion of augmented or synthetic images 352 in the preprocessed images 340. For example, it may be beneficial to limit the number or proportion of synthetic images 352 included in the subset of training images 340a. In some examples, the ratio of original image 335 to composite image 352 is maintained at 1:1, 1:2, 2:1, 1:3, 3:1, 1:4, or 4:1.

[0083] In some examples, the machine learning algorithm 355 includes a CNN, a modified CNN with the encoding layer replaced by a residual neural network ("Resnet"), or a modified CNN with the encoding and decoding layers replaced by a Resnet. In other examples, the machine learning algorithm 355 may be any suitable machine learning algorithm configured to localize, classify, and / or analyze the preprocessed image 340, such as a two-dimensional CNN ("2DCNN"), Mask R-CNN, U-Net, feature pyramid network (FPN), dynamic time warping ("DTW") technique, hidden Markov model ("HMM"), pure attention-based model, or one or more combinations of such techniques, such as a visual transformer, CNN-HMM, or MCNN (multiscale convolutional neural network). The computing environment 300 may employ the same type of machine learning algorithm or different types of machine learning algorithms trained to detect and classify different cells. For example, the computing environment 300 may include a first machine learning algorithm (e.g., U-Net) for detecting and classifying PD1. The computing environment 500 may also include a second machine learning algorithm (e.g., 2DCNN) for detecting and classifying cluster of differentiation 68 ("CD68"). The computing environment 300 may also include a third machine learning algorithm (e.g., U-Net) for detecting and classifying PD1 and CD68 in combination. The computing environment 300 may also include a fourth machine learning algorithm (e.g., HMM) for diagnosing disease for treatment or prognosis of a subject, such as a patient. In other examples according to the present disclosure, still other types of machine learning algorithms may be implemented.

[0084] The training process of the machine learning algorithm 355 includes selecting hyperparameters for the machine learning algorithm 355 from the parameter data store 363, inputting a subset of images 340a (e.g., labels 345 and corresponding preprocessed images 340) into the machine learning algorithm 355, and performing iterative operations to learn a set of parameters (e.g., one or more coefficients and / or weights) for the machine learning algorithm 355. Hyperparameters are settings that can be adjusted or optimized to control the behavior of the machine learning algorithm 355. Most algorithms explicitly define hyperparameters that control different aspects of the algorithm, such as memory or execution cost. However, additional hyperparameters may be defined to adapt the algorithm to specific scenarios. For example, hyperparameters may include the number of hidden units of the algorithm, the learning rate of the algorithm (e.g., 1e-4), the convolution kernel width, or the number of kernels of the algorithm. In some examples, the number of model parameters is reduced for each convolutional and deconvolutional layer, and / or the number of kernels is reduced by half for each convolutional and deconvolutional layer compared to a typical CNN.

[0085] The subset of images 340a may be input to the machine learning algorithm 355 in batches of a predetermined size. The batch size limits the number of images shown to the machine learning algorithm 355 before parameter updates can be performed. Alternatively, the subset of images 340a may be input to the machine learning algorithm 355 as a time series or sequentially. In either case, if augmented or composite images 352 are included in the preprocessed images 340a, the number of original images 335 versus the number of composite images 352 included in each batch, or the manner in which the original images 335 and phenotypic images 352 are fed to the algorithm (e.g., every other batch or image is an original batch of images or an original image), can be defined as a hyperparameter.

[0086] Each parameter is a tunable variable such that values ​​for the parameter are adjusted during training. For example, a cost function or objective function may be configured to optimize accurate classification of depicted expressions, optimize characterization of a given type of feature (e.g., characterization of shape, size, uniformity, etc.), optimize detection of a given type of feature, and / or optimize accurate localization of a given type of feature. Each iteration may include learning a set of parameters for the machine learning algorithm 355 that minimizes or maximizes a cost function of the machine learning algorithm 355, such that the value of the cost function using a set of parameters may be smaller or larger than the value of the cost function using a different set of parameters in a previous iteration. The cost function may be configured to measure the difference between the predicted output using the machine learning algorithm 355 and the labels 345 included in the training data. For example, in the case of a model based on supervised learning, the goal of training is to learn a function “h()” (sometimes called a hypothesis function) that maps a training input space X to a target value space Y, h: X → Y, where h(x) is a good predictor for the corresponding value of y. A variety of different techniques may be used to learn this hypothesis function. In some techniques, as part of deriving the hypothesis function, a cost or loss function may be defined that measures the difference between a ground truth value for an input and a predicted value for that input. As part of training, techniques such as backpropagation, random feedback, direct feedback alignment (DFA), indirect feedback alignment (IFA), and Hebbian learning are used to minimize this cost or loss function.

[0087] The training iterations continue until a stopping condition is met. The training completion condition may be configured to be met (for example) when a predetermined number of training iterations are completed, when a statistical value generated based on testing or validation exceeds a predetermined threshold (e.g., a classification accuracy threshold), when a statistical value generated based on a confidence metric (e.g., the mean or median of the confidence metric or the percentage of the confidence metric above a particular value) exceeds a predetermined confidence threshold, and / or when a user device involved in the training review closes the training application executed by the training controller 365. Once a set of model parameters is identified through training, the machine learning algorithm 355 is trained, and the training controller 365 performs an additional process of testing or validation using the subset of images 340b (the test or validation dataset). The validation process may include iterative operations of inputting images from the subset of images 340b into the machine learning algorithm 355 using validation techniques such as k-fold cross-validation, leave-one-out cross-validation, leave-one-out cross-validation, and nested cross-validation to adjust the hyperparameters and ultimately find an optimal set of hyperparameters. Once an optimal set of hyperparameters is obtained, a reserved test set of images from the subset of images 340b is input into the machine learning algorithm 355 to obtain outputs, which are evaluated against the ground truth using correlation techniques such as Bland-Altman or Spearman's rank correlation coefficient to calculate performance metrics such as error, accuracy, precision, recall, and receiver operating characteristic curves (ROC). In some examples, a new training iteration may be initiated in response to receiving a corresponding request or trigger condition from a user device (e.g., initial model development, model update / adaptation, continuous learning, drift determined within the trained machine learning model 360, etc.).

[0088] As will be appreciated, other training / validation mechanisms are contemplated and may be implemented within computing environment 300. For example, machine learning algorithm 355 may be trained and hyperparameters may be tuned on images from subset of images 340a, while images from subset of images 340b may be used solely to test and evaluate the performance of machine learning algorithm 355. Furthermore, the training mechanisms described herein focus on training new machine learning algorithms 355. These training mechanisms may also be utilized for initial model development, model updating / adaptation, and continuous learning of existing machine learning models 360 trained from other datasets, as described in detail herein. For example, in some cases, machine learning model 360 may have been preconditioned using images of other objects or biological structures, or from sections from other subjects or studies (e.g., human or mouse studies). In those cases, machine learning model 360 may be used for initial model development, model updating / adaptation, and continuous learning using preprocessed images 340.

[0089] The trained machine learning model 360 can then be used (in the result generation stage 325) to process the new preprocessed image 340 to generate predictions or inferences, such as predicting cell center and / or location probabilities, classifying cell types, generating cell masks (e.g., segmentation masks for each pixel of the image), predicting a disease diagnosis or prognosis for a subject, such as a patient, or any combination thereof. In some examples, the masks identify the locations of depicted cells associated with one or more biomarkers. For example, given tissue stained for a single biomarker, the trained machine learning model 360 can be configured to (i) infer cell centers and / or locations, (ii) classify cells based on features of the staining pattern associated with the biomarker, and (iii) output a cell detection mask for positive cells and a cell detection mask for negative cells. As another example, given tissue stained for two biomarkers, the trained machine learning model 360 may be configured to (i) infer cell centers and / or locations, (ii) classify the cells based on features of the staining patterns associated with the two biomarkers, and (iii) output a cell detection mask for cells positive for the first biomarker, a cell detection mask for cells negative for the first biomarker, a cell detection mask for cells positive for the second biomarker, and a cell detection mask for cells negative for the second biomarker. As another example, given tissue stained for a single biomarker, the trained machine learning model 360 may be configured to (i) infer cell centers and / or locations, (ii) classify the cells based on features of the cells and the staining patterns associated with the biomarkers, and (iii) output a cell detection mask for positive cells and a cell detection mask for negative cell codes, as well as mask cells classified as tissue cells.

[0090] In some examples, analysis controller 380 generates analysis results 385 that are utilized by the entity that requested processing of the underlying image. Analysis results 385 may include a mask output from trained machine learning model 360 overlaid on new preprocessed image 340. Additionally or alternatively, analysis results 385 may include information calculated or determined from the output of the trained machine learning model, such as a whole-slide tumor score. In exemplary embodiments, the automated analysis of the tissue slide uses assignee VENTANA's FDA-cleared 510(k)-approved algorithm. Alternatively or additionally, any other automated algorithm may be used to analyze selected regions of the image (e.g., masked images) and generate a score. In some embodiments, analysis controller 380 may further respond to instructions received from a computing device from a pathologist, physician, investigator (e.g., associated with a clinical trial), subject, medical professional, etc. In some examples, the communication from the computing device includes an identifier for each of a particular set of subjects and corresponds to a request to perform an iteration of the analysis for each subject represented in the set. The computing device may further perform analysis based on the machine learning model and / or the output of analysis controller 380 and / or may provide a recommended diagnosis / treatment to the subject.

[0091] It will be understood that computing environment 300 is exemplary and that computing environments 300 having different stages and / or using different components are contemplated. For example, in some instances, the network may omit pre-processing stage 310, such that the images used to train the algorithm and / or images processed by the model are raw images (e.g., from an image data store). As another example, it will be understood that each of pre-processing stage 310 and training stage 320 can include a controller for performing one or more operations described herein. Similarly, although labeling stage 315 is depicted in relation to labeling controller 350 and result generation stage 325 is depicted in relation to analysis controller 380, the controllers associated with each stage may additionally or alternatively facilitate other operations described herein other than generating labels and / or generating analysis results. 3 lacks depicted representations of devices associated with a programmer (e.g., who selected the architecture of the machine learning algorithm 355, which defined how various interfaces function, etc.), devices associated with a user providing an initial label or label review (e.g., in the labeling stage 315), and devices associated with a user requesting model processing for a given image (which may be the same user or a different user than the user who provided the initial label or label review). Despite the lack of depictions of these devices, the computing environment 300 may include use of one, more, or all of the devices, and may in fact include use of multiple devices associated with a corresponding number of users providing initial labels or label reviews and / or a corresponding number of users requesting model processing for various images.

[0092] V. Determining the consensus position of annotations When annotations from different annotators are added to an image, various types of annotation conflicts may inevitably arise. For example, one or more annotation location conflicts between annotations may be found in an image. An annotation-location conflict may be identified when (i) one annotation identifies a cell located at a specific location in the image, but one or more other annotations identify the same cell as located at a different location in the image; (ii) one annotation identifies multiple cells located at a given location, but one or more other annotations identify a single cell located at the same given location; or (iii) one annotation identifies a cell located at a location, but one or more other annotations identify no cell located at that location. Thus, determining a consensus location between annotations ensures that corresponding labels point to the same location and can remove annotations (e.g., duplicate annotations) that may not accurately identify cells in the image.

[0093] A. Input Data To receive and process annotations associated with a digital pathology image (or the FOV of an image), an annotation processing system (e.g., annotation processing system 102 of FIG. 1 , annotation processing system 351 of FIG. 3 ) parses one or more schema files (or other types of data structures) provided by a set of annotators. The schema files can include a set of annotations, each identifying an initial location of a particular image region in the image and an initial label (e.g., stained tumor cell) that classifies the biological structure depicted in the initial location on the image. The annotations can be added by the annotators using an annotation tool (e.g., dpath). A set of protocols can be defined for the annotators so that annotations can be generated based on configurations for the same image. The instructions can include pre-configured values ​​for: (i) magnification level, (ii) cell type to annotate, (iii) definition of each label, (iv) additional indicators in case of ambiguity, (v) use of specific color profiles, (vi) seed placement of annotations, and (iv) showing cells at the border of the image.

[0094] In addition to the above information, the annotation can be associated with a location within the image. In some examples, the annotation includes x and y coordinates that identify its location relative to the image. Additional data can be added to the annotation based on input provided by the annotator. The annotations from the annotator can be exported as a schema file, such as an XML file. In some examples, a single schema file is generated for the entire image. Additionally or alternatively, a schema file can be generated for each FOV of the image.

[0095] The annotation processing system can receive and parse the schema file to determine the corresponding image (or FOV of the image) to which annotations should be added and associate each annotation with a corresponding position on the image. The annotated images can be processed as part of a training dataset used to train a machine learning model. In some examples, the images are processed to be part of a validation dataset used to validate the machine learning model. Additionally or alternatively, the images can be processed to be part of a test dataset used to test the machine learning model.

[0096] B. Identifying annotation clusters As the annotations are processed, the annotation processing system determines which annotations submitted by different annotators indicate the same location. Figure 4 shows a schematic diagram illustrating a process 400 for identifying clusters for a set of annotations of a slide image, according to some embodiments. Process 400 iterates through the annotations submitted by each annotator to identify clusters, thereby enabling annotations that correspond to the same region to be associated.

[0097] 4, the plurality of annotations associated with the image may include a first set of annotations 402 associated with a first annotator, "Path 1." Each annotation in the first set 402 includes an identifier (e.g., "p11") and a set of x and y coordinates that identify the annotation's location within the image. The annotations in the first set of annotations 402 may be compared with annotations from a set of remaining annotations 404. The set of remaining annotations 404 may include a second set of annotations associated with a second annotator, "Path 2," and a third set of annotations associated with a third annotator, "Path 3," where the annotations in the second and third sets include a respective identifier (e.g., "p21") and a respective set of x and y coordinates.

[0098] Based on the above information, the annotation processing system can perform clustering of the multiple annotations based on their respective x and y coordinates, thereby forming one or more annotation clusters. In some examples, the one or more clusters are formed by performing a distance-based search among the multiple annotations. A cluster can be defined as a set of annotations that are within a predetermined distance. In some examples, a threshold is used to determine the set of annotations, the threshold identifying a number of pixels (e.g., 10 pixels) between the annotations. The threshold can be adjusted by the annotator, which can increase or decrease the number of annotations included in the cluster. Adjusting the threshold to increase or decrease the number of annotations can define the size of the cluster for the image.

[0099] As shown in FIG. 4 , the first set of clusters 406 includes a first cluster (shown in pink) that includes the annotation “p11” from the first set of annotations 404, as well as the annotations “p22” and “p31.” The annotation “p11” may include identifiers of “p22” and “p31” as nearest points (“Cp-id”). The first cluster may then be identified by the association key of “p11” for the first cluster of annotations “p11,” “p22,” and “p31.” Similarly, the first set of clusters 406 may include a second cluster (shown in blue) that includes the annotations “p12,” “p21,” and “p32,” and the second cluster may be identified by its association key “p12.”

[0100] Additional cluster combinations can be formed based on repeating the above process through other sets of annotations (e.g., second and third sets of annotations). For example, the second set of annotations 408 associated with "Pathway 2" of a second annotator can be compared with another set of remaining annotations 410. The other set of remaining annotations includes a first set of annotations corresponding to "Pathway 1" and a third set of annotations corresponding to "Pathway 3." Clustering can be performed to generate a cluster of second annotations 412, which can include annotation "p23" of the second set of annotations 408 associated with "p32" of the third set of annotations. Furthermore, "p22" and "p31" are associated with the same cluster based on the associated key "p11." In some examples, clustering is iterated through each set of annotations until all possible combinations of annotation clusters have been formed.

[0101] The first cluster 406 and the second cluster 412 may then be merged to generate a merged cluster of annotations 414. The merged cluster may be generated by merging a cluster from the first set of clusters 406 with a corresponding cluster from the second set of clusters 412. For example, the annotation "p32" may be identified as the closest annotation to the annotation "p12" from the first set of annotations 404 and the annotation "p23" from the second set of annotations 408. Based on this association, the "p23" annotation may be merged into the cluster with the related key "p12." In another example, the annotation "p33" may be identified as the closest annotation to the annotation "p21," which may also be identified as the closest annotation to the annotation "p12." As a result, the annotation "p33" may also be merged into the cluster with the related key "p12."

[0102] FIG. 5 illustrates an exemplary image 500 including annotation locations, according to some embodiments. Each depicted cell in image 500 is associated with one or more colored circles (e.g., pink, orange, blue, green, red), each representing an annotation 502. The color of the circle identifies the annotator who generated the annotation. For example, a red circle identifies a first pathologist who submitted an annotation for image 500, and a blue circle identifies a second pathologist who submitted an annotation for image 500. Annotations that are within a predetermined distance threshold (e.g., 10 pixels) can be associated as part of a cluster 504. A cluster 504 can represent a single cell and can be identified by a yellow or black line connecting the circles. In some examples, two or more clusters are merged into a single merged cluster 506.

[0103] C. Removing annotation conflicts When clusters are generated for annotations, annotation location conflicts can arise when two or more annotations (representing multiple cells) from a particular annotator are identified during distance-based search as referring to the same cluster (representing a single cell). Annotation location conflicts can arise when two cells are located very close to each other in a given image. In such cases, it becomes difficult to determine which cell a particular annotation represents. In another example, a pathologist may erroneously indicate the presence of two tumor cells for a single tumor cell, thereby causing an annotation location conflict.

[0104] 6 shows an example set of annotations identified in a training slide image 600, according to some embodiments. Each depicted cell in image 600 is associated with one or more colored circles (e.g., pink, orange, blue, green, red), each representing an annotation. The color of the circle identifies the annotator who submitted the annotation. The circles in image 600 are connected by black or yellow lines, which represent particular clusters. A cluster can represent or identify a single cell.

[0105] In a single cluster, two or more annotations (e.g., multiple cells) can be identified from the same annotator. For example, cluster 602 can be identified by at least three orange circles, and the corresponding annotator considers the image region to depict three cells. Because cluster 602 is presumed to represent only a single cell but now contains annotations indicating multiple cells, based on the three annotations from the same annotator, the annotation processing system may determine that cluster 602 contains at least one annotation position conflict. Annotation-position conflicts can be indicated by the annotation processing system by tagging cluster 602 (e.g., the cluster's associated key) as having a "conflict" and displaying such clusters in a graphical user interface.

[0106] In some examples, the annotation processing system indicates annotation location conflicts by providing visual markers in the image 600. For example, the first set of clusters 604 may be indicated as having no annotation location conflicts based on one or more "black" lines connecting the annotations. In contrast, the second set of clusters 606 may be indicated as having one or more annotation location conflicts based on one or more "yellow" lines connecting the annotations.

[0107] Two types of annotation-location conflicts may be encountered during processing, which may include: (i) a first conflict type involving a location with two or more annotations from a single annotator, and (ii) a second conflict type involving a location with two or more annotations from each of two or more annotators.

[0108] FIG. 7 shows an example image illustrating a process 700 for resolving annotation conflicts, according to some embodiments. Process 700 can include removing annotation conflicts for a first conflict type. Pre-processed image 702 shows an example image including one or more annotation conflicts, including a cluster 704. The annotation processing system can determine, for each of two or more annotations from the same annotator, a statistical value (e.g., mean, median) representing the distance from the annotation to other annotations in the corresponding cluster. The statistical values ​​between two or more annotations from the same annotator can be compared. The annotation with the lowest statistical value can be determined to be part of the cluster, while the remaining annotations in the conflict can be removed from being part of the cluster. Post-processed image 706 shows another example image in which one or more annotation conflicts have been removed. As shown in post-processed image 706, the yellow line connecting orange circle 708 to cluster 704 has been removed. In some examples, the orange circle 708 may form its own cluster and identify different objects or regions within the image.

[0109] FIG. 8 shows another example set of images illustrating a process 800 by which annotation conflicts are removed, according to some embodiments. A preprocessed image 802 shows an example image including one or more annotation conflicts, including a merged cluster 804. The process 800 can include removing annotation conflicts for a second conflict type. The annotation processing system can apply a clustering algorithm to a cluster including two or more annotations from each of two or more annotators, such that the cluster can be divided into sub-clusters. In some examples, the clustering algorithm includes a k-means algorithm that divides the cluster into "k" sub-clusters. The "k" number of sub-clusters can correspond to the maximum number of annotations entered by a particular annotator for a given cluster.

[0110] For each sub-cluster containing two or more annotations from the same annotator, the annotation processing system may determine, for each of the two or more annotations, a statistical value (e.g., mean, median) that represents the distance from the annotation to other annotations in the corresponding cluster. The statistical values ​​between two or more annotations from the same annotator may be compared. The annotation with the lowest statistical value may be determined to be part of the sub-cluster, while the remaining annotations may be excluded from being part of the sub-cluster. Post-processed image 806 shows another exemplary image in which one or more annotation conflicts have been removed. Sub-clusters 808 and 810 show two clusters separated from merged cluster 804, with the yellow line connecting sub-clusters 808 and 810 removed.

[0111] D. Determine the consensus position After the annotation conflicts are resolved, the annotation processing system can determine a set of clusters that can remain in the image. A cluster in the set of clusters can include some annotations that exceed a confidence threshold. The confidence threshold can correspond to a minimum number of annotations that must be present within a given cluster. For example, the confidence threshold can specify that at least three annotations must be present for a cluster to remain in the image. A cluster that includes four annotations can be determined to exceed the confidence threshold and can be added to the set of clusters that can remain in the image. Conversely, a cluster that includes two annotations can be determined to not exceed the confidence threshold. Such a cluster is discarded from the image, and the cluster's annotations are also discarded from the image. As a result, the set of clusters with a sufficient number of annotations can then be further processed. In some examples, the confidence threshold is a percentage value (e.g., 0.5, 40%) compared to a statistical value representing the proportion of annotations present in the cluster to the total number of annotators who generated annotations for the image. The confidence threshold can be adjusted based on user input, which can result in an increase or decrease in the ground truth labels that will be present in the image.

[0112] After determining clusters that exceed a confidence threshold, a consensus position can be determined for each cluster identified in the image. Each cluster can include a set of annotations submitted by multiple annotators, and the consensus position of the cluster can be determined based on the median coordinate between the positions of the set of annotations associated with the cluster. For example, the x- and y-coordinates of each annotation in a given cluster can be identified. The median x-coordinate can be determined from the x-coordinates of the set of annotations for the cluster. Similarly, the median y-coordinate can be determined from the y-coordinates of the set of annotations for the cluster. The median x-coordinate and the median y-coordinate can be used to determine the consensus position of the cluster so that the ground truth label can be placed at the consensus position of the image. Additionally or alternatively, other statistical values, such as the mean, can be used instead of the median x- and y-coordinates.

[0113] In some examples, the x and y coordinates of each annotation are assigned a corresponding weighted value based on the annotator who submitted the annotation. For example, the x and y coordinates of an annotation submitted by a senior pathologist may be assigned a weighted value greater than another weighted value assigned to the x and y coordinates of another annotation submitted by a junior pathologist. Thus, a consensus position may be determined based on the weighted x and y coordinates of a set of annotations within a given cluster. The consensus position of a cluster may be assigned a ground truth label that identifies the target type of biological structure (e.g., tumor cell) of the corresponding object or region in the image. The ground truth label at the consensus position may also be used to train a machine learning model that predicts the type of biological structure in the corresponding region of other slide images.

[0114] E. Methods for Identifying Annotation-Location Conflicts in Images 9 shows a process 900 for identifying image annotation-location conflicts, according to some embodiments. For illustrative purposes, the process 900 is described with reference to the components shown in FIG. 1 and / or FIG. 3, although other implementations are possible. For example, program code for the annotation processing system 102 of FIG. 1 and / or the annotation processing system 351 of FIG. 3 stored on a non-transitory computer-readable medium may be executed by one or more processing devices to cause a server system to perform one or more operations described herein.

[0115] In step 902, the annotation processing system accesses a plurality of annotations associated with an image depicting at least a portion of a biological sample. The image may be a digital pathology image depicting at least a portion of a biological sample (e.g., tissue) obtained from a subject. The plurality of annotations can be used to generate ground truth labels for supervised learning, validation, and / or testing of a machine learning model. Each annotation of the plurality of annotations can be located within the image at a position that identifies an object or region within the image. In some examples, each annotation includes an identifier associated with the annotator (e.g., a pathologist) who generated the annotation.

[0116] In step 904, the annotation processing system identifies, from the plurality of annotations, a first annotation generated by the first annotator and a second annotation generated by the second annotator. In some examples, the first annotation further includes a label classifying a target type of a biological structure depicted at the first image location (e.g., stained tumor cell, unstained tumor cell, normal cell). Similarly, the second annotation includes a respective label classifying the same or a different target type of a biological structure depicted at the second image location.

[0117] In step 906, the annotation processing system determines a first distance value between the first position of the first annotation and the second position of the second annotation.

[0118] In step 908, the annotation processing system determines that the first distance value is below a predetermined threshold. The predetermined threshold can be used to determine whether the first and second annotations correspond to the same object or region in the image. In some examples, the predetermined threshold is adjusted to increase or decrease the number of annotations that collectively identify the same object or region of the image.

[0119] In step 910, the annotation processing system determines that the first annotation and the second annotation both identify a first object or region in the image in response to determining that the first distance value is below a predetermined threshold.

[0120] In step 912, the annotation processing system identifies a third annotation generated by the first annotator from the plurality of annotations, and the third annotation is placed in a third location different from the first location.

[0121] In step 914, the annotation processing system determines that the value of the second distance between the second position of the second annotation and the third position of the third annotation is below a predetermined threshold.

[0122] In step 916, the annotation processing system determines a third annotation to also identify the first object or region in the image in response to determining that the second distance value is below the predetermined threshold.

[0123] In step 918, in response to determining the third annotation as identifying the first object or region in the image, the annotation processing system generates an output indicating that the object or region in the image has an annotation conflict, with process 900 then terminating.

[0124] F. Image Annotation - Methods for Removing Location Conflicts 10 shows a process 1000 for removing annotation-location conflicts and determining a consensus label for an image, according to some embodiments. For illustrative purposes, the process 1000 is described with reference to the components shown in FIG. 1 and / or FIG. 3, although other implementations are possible. For example, program code for the annotation processing system 102 of FIG. 1 and / or the annotation processing system 351 of FIG. 3 stored on a non-transitory computer-readable medium may be executed by one or more processing devices to cause a server system to perform one or more operations described herein.

[0125] In step 1002, the annotation processing system identifies positional conflicts among a first annotation, a second annotation, and a third annotation of an image. The training image may be a digital pathology image depicting at least a portion of a biological sample (e.g., tissue) obtained from a subject. The first, second, and third annotations can be used to generate ground truth labels for supervised learning, validation, and / or testing of a machine learning model. The first and third annotations were generated by a first annotator, and the second annotation was generated by a second annotator. The first, second, and third annotations can be identified as identifying a first object or region in the image. Annotation conflicts can be determined using the steps of process 900 of FIG. 9.

[0126] In step 1004, the annotation processing system compares a first distance value between the first annotation and the second annotation with a second distance value between the second annotation and the third annotation. The first distance value may correspond to a distance between a first position of the first annotation and a second position of the second annotation. The second distance value may correspond to a distance between a second position of the third annotation and a third position.

[0127] In step 1008, if the first distance value is greater than the second distance value (the "Yes" path from step 1006), the annotation processing system may resolve the annotation-location conflict by determining that the third annotation corresponds to a second object or region of the image (e.g., a different cluster). In step 1010, if the second distance value is greater than the first distance value (the "No" path from step 1006), the annotation processing system may resolve the annotation location conflict by determining that the first annotation corresponds to a second object or region of the image (e.g., a different cluster).

[0128] In step 1012, the annotation processing system determines a consensus location of the first object or region of the image based on the remaining annotations associated with the first object or region. The consensus location can be determined based on a median coordinate between the locations of the remaining annotations associated with the first object or region of the image. Additionally or alternatively, other statistical values, such as a mean value, can be used instead of the median x and y coordinates.

[0129] In some examples, the annotation processing system resolves another type of annotation location conflict involving two or more annotations from two or more annotators. For example, the annotation processing system identifies a fourth annotation generated by a second annotator. The fourth annotation is located at a fourth image location different from the second location of the second annotation. The annotation processing system then determines a third distance value between the third location and the fourth location and further determines that the third distance value is below a predetermined threshold. In response to determining that the third distance value is below the predetermined threshold, the annotation processing system determines that the first annotation, the second annotation, the third annotation, and the fourth annotation identify the same object or region in the image. Because two or more annotations from each of the first and second annotators have been identified for the same region, the annotation processing system can determine that a different type of annotation location conflict has occurred.

[0130] To resolve other annotation-location conflicts, the annotation processing system can apply a clustering algorithm to the first location, the second location, the third location, and the fourth location. The clustering algorithm can be applied to determine that (i) the first annotation and the second annotation correspond to a first sub-region of the first object or region, and (ii) the third annotation and the fourth annotation correspond to a second sub-region of the first object or region. In some cases, the clustering algorithm is a k-means clustering algorithm.

[0131] In step 1014, the annotation processing system assigns a ground truth label for the consensus location of the image. The ground truth label identifies the target type (e.g., tumor cell) of the biological structure of the first object or region in the image. The ground truth label at the consensus location can be used to train a machine learning model that predicts the type of biological structure in the corresponding region of the other slide image. Training the machine learning model using the ground truth label with the consensus location can correspond to the training process described in FIG. 3. For different types of annotation-location conflicts that resulted in the first and second subregions, the annotation processing system can assign respective ground truth labels for each of the first and second subregions. Process 1000 then ends.

[0132] VI. Determining Consensus Labels for Annotations After a consensus location is determined for a set of annotations, the annotation processing system can determine a ground truth label for the consensus location on the image, which can be used to train, validate, and / or test a machine learning model. In some examples, a ground truth label is determined when all of the set of annotations indicate the same biological structure target type (e.g., tumor cells). Similarly, a ground truth label can be determined based on a set of annotations that identify a clear majority of a particular biological structure target type. However, there are instances where there is no clear consensus among the sets of annotations regarding the biological structure target type. For example, a consensus location may include four annotations, where two annotators label the location as having stained tumor cells, while two other annotators label the location as having unstained tumor cells.

[0133] In the case of such annotation-label conflicts, the annotation processing system can determine a consensus label for an object or region of the image. In particular, each annotation in the set of annotations can be assigned a corresponding weight value. The weight value is determined based on performance and historical data associated with the annotator who generated the annotation. For example, the annotator of a first annotation can be a senior pathologist with many years of experience. The weight of the first annotation can be assigned a relatively higher value than the weight of another, second annotation associated with a junior pathologist.

[0134] Each target type of biological structure can then be represented by an aggregate weight derived from the weights assigned to the corresponding annotations. In response to determining that a particular target type of biological structure has the highest aggregate weight, a ground truth label can be generated to include the particular target type of biological structure. As a result, a consensus label can be determined for regions or objects in the image whose ground truth label can include information from the consensus label. Determining a consensus label can facilitate accurate training, validation, and testing of machine learning models, even when several annotations provide conflicting information at the same location.

[0135] A. Determine whether a set of annotations can be used to determine a consensus label for a region of the image The annotation processing system can determine whether a set of annotations corresponding to a particular region or object (e.g., consensus location) of an image can be used to generate a consensus label.

[0136] FIG. 11 shows a schematic diagram illustrating a process 1100 for removing annotation-label conflicts and determining a consensus label for an image, according to some embodiments. At block 1102, an annotation processing system receives a set of annotations corresponding to regions or objects of an image. The images may be part of a training dataset used to train a machine learning model. In some examples, the images are part of a validation dataset used to validate a machine learning model. Additionally or alternatively, the images are part of a test dataset used to test a machine learning model. The regions or objects may be identified based on the consensus positions determined using the processes described in FIGS. 9 and 10.

[0137] In block 1104, the annotation processing system may obtain weights associated with the set of annotations. For each annotation in the set of annotations, a weight may be assigned to the annotation, and the weight may be determined based on performance and historical data associated with the annotator who generated the annotation. For example, the weight may be determined based on the level of expertise of the annotator who generated the annotation. In some examples, each weight in the weights corresponds to a value between 0 and 1, and the sum of the weights of all annotators may add up to 1.

[0138] In some examples, the annotation processing system calculates a sum of the weights and compares the sum to a predetermined confidence threshold. The confidence threshold is a user-adjustable value that can specify the number of annotations required to generate a consensus label. Comparing the weights to the confidence threshold can facilitate determining whether a set of annotations contains sufficient information to generate a consensus label.

[0139] As an illustrative example, two annotations corresponding to a region of an image may be determined to be sufficient to generate a consensus label based on the sum of their corresponding weights exceeding a confidence threshold. In another example, four annotations corresponding to another region of an image may be determined to be insufficient to generate a consensus label based on the sum of their corresponding weights not exceeding a confidence threshold. In some examples, the annotation processing system includes one or more other annotations corresponding to locations proximate to a particular region or object so that some annotations with higher weights can also be considered.

[0140] In block 1106, the annotation processing system may determine whether the set of annotations is sufficient to generate a consensus label based on comparing the count of the set of annotations for a particular region or object with a second confidence threshold. For example, the second confidence threshold may be three, and if the number is three or greater, the set of annotations may be deemed sufficient to generate a consensus label. If the count of the set of annotations does not exceed the second confidence threshold, the annotation processing system may discard the set of annotations from the particular region or object so that the set of annotations is no longer processed to generate a ground truth label. In contrast, if the count of the set of annotations exceeds the second confidence threshold, the annotation processing system may use the set of annotations to generate a consensus label. This process may be repeated for annotations from each of the other regions or objects of the image. As a result, sets of annotations across different regions or objects may then be used to generate a consensus label for the image.

[0141] B. Resolving label conflicts in a set of annotations For a set of annotations corresponding to an object or region of an image, the annotation processing system can eliminate any conflicts in the labels and generate a consensus label for the object or region. Referring back to FIG. 11 , the annotation processing system can identify a target type of biological structure based on the weights assigned to the set of annotations (block 1108). For each target type of biological structure identified in the set of annotations, an aggregated weight can be determined based on the subset of annotations that includes the target type of biological structure. For example, a first aggregated weight for a target stained tumor cell can be based on the weights of three annotations in the set of annotations that identify the corresponding region of the image that depicts tumor cells. In another example, a second aggregated weight for a target immune cell (e.g., lymphocyte) can be based on the weights of the two remaining annotations in the set of annotations that identify the corresponding region of the image that depicts immune cells. In some examples, the aggregated weight corresponds to the sum of all weights associated with the corresponding target type of biological structure. Additionally or alternatively, the aggregated weight can correspond to the mean or median of all weights associated with the corresponding target type of biological structure.

[0142] The annotation processing system can compare the aggregated weights to determine the target type of the biological structure with the highest weight. The target type of the biological structure with the highest aggregated weight can be identified as the consensus label for the region or object of the image. As a result, a ground truth label including the target type can then be generated for the image. Continuing with the above example, it can be determined that the second aggregated weight for the targeted immune cell is greater than the first aggregated weight, even if the targeted stained tumor cell is associated with a greater number of annotations. In response to the determination, the annotation processing system generates a ground truth label for the corresponding region of the image whose ground truth label includes the immune cell.

[0143] 12 shows an example image 1200 including consensus labels and locations, according to some embodiments. Each dot represents a consensus location depicting a biological structure (e.g., tumor cell). The color of the dot represents the ground truth label of the location. For example, "blue" dots represent other types of cells (e.g., lymphocytes, fibroblasts), "red" dots represent stained tumor cells, and "green" dots represent unstained tumor cells.

[0144] In some examples, there are situations in which at least two aggregated weights for corresponding regions of an image are equal to each other. Referring back to FIG. 11 , the annotation processing system may determine a quadratic weight for each annotation in the set of annotations (block 1110). The quadratic weight may correspond to a value assigned based on an accuracy metric assigned to the annotator who generated the annotation. The accuracy metric may be determined based on several other annotations provided by the annotator for other regions of the image that match the respective consensus label. The quadratic weights may be used to normalize the aggregated weights, such that a target type of biological structure associated with a larger normalized weight may be identified as the consensus label for the corresponding region.

[0145] As an illustrative example, an object or region of an image includes two annotations, where a first annotation indicates a stained tumor cell and a second annotation indicates an unstained tumor cell. The annotation processing system determines that the two annotations have equal weights assigned to them, thereby identifying a label conflict between the annotations. In response to such a determination, a secondary weight for the first annotation may be calculated based on a determination that 90% of the annotations generated by an annotator for the other region are consistent with the corresponding consensus label. Then, a secondary weight for the second annotation may be calculated based on a determination that 80% of the annotations generated by another annotator for the other region are consistent with the corresponding consensus label, where the secondary weight of the second annotation is less than the secondary weight of the first annotation based on their respective accuracy metrics. Based on the comparison, the stained tumor cell indicated by the first annotation may be selected to generate a ground truth label for the object or region of the image. Thus, by using the weights and quadratic weights, the annotation processing system can resolve conflicts between annotation labels in different regions of an image.

[0146] C. Adjust weights based on annotator performance The consensus labels can be used to train machine learning models, but also to reevaluate annotators and adjust their corresponding weights. Referring to FIG. 11 , the annotation processing system can adjust the weights associated with an annotator by tracking its performance across different images (e.g., training images, validation images) or different FOVs within the same image (block 1112). For an annotator that generated one or more annotations for an image (or FOV of an image), the annotation processing system can determine the number of annotations that match the respective consensus label for the given image. The number of annotations can be compared to the total number of annotations generated by the annotator for the image. Based on the comparison, the annotation processing system can adjust the weights associated with the annotator. The weight adjustment process can continue across other images to fine-tune the annotator's weights. The above process can be repeated across other annotators that generated annotations for the image. Additionally or alternatively, the annotation processing system may adjust the annotation weights based on the number of annotators that agree with each consensus label for each target type of biological structure defined for the image.

[0147] In some examples, an annotator's weight is adjusted based on its performance relative to the performance of other annotators for an image. For example, a first annotator may be determined to have 85% of annotations that match the image's respective consensus label. A second annotator may be determined to have 89% of annotations that match the image's respective consensus label. The performance of the first and second annotators may be compared, and based on the comparison, the second annotator may be assigned a higher weight (e.g., a weight assigned proportional to the percentage of annotations that match the respective consensus labels) than that of the first annotator. If three or more annotators generated annotations for an image, the annotators may be ranked based on their performance for the image, and the corresponding weights may be adjusted based on the ranking.

[0148] In practice, when the fine-tuning step is performed across multiple FOVs of images (e.g., 100 FOVs) or multiple images, the weights can accurately reflect the annotator's expertise and performance. In some examples, particular annotators assigned low weights are retrained before being allowed to provide annotations to other images. Additionally or alternatively, annotations of particular annotators assigned low weights may be excluded from use during the process for generating ground truth labels for images.

[0149] D. Image Annotation - Methods for Removing Label Conflicts 13 illustrates a process 1300 for removing annotation-label conflicts and determining a consensus label for an image, according to some embodiments. For illustrative purposes, the process 1300 is described with reference to the components shown in FIG. 1 and / or FIG. 3, although other implementations are possible. For example, program code for the annotation processing system 102 of FIG. 1 and / or the annotation processing system 351 of FIG. 3 stored on a non-transitory computer-readable medium may be executed by one or more processing devices to cause a server system to perform one or more operations described herein.

[0150] In step 1302, the annotation processing system receives a plurality of annotations corresponding to a first object or region of an image. The training images may be digital pathology images depicting at least a portion of a biological sample (e.g., tissue) obtained from a subject. The plurality of annotations can be used to generate ground truth labels for supervised learning, validation, and / or testing of a machine learning model. Each annotation in the set of annotations identifies a target type of a biological structure depicted in the first object or region. For example, target types of biological structures include tumor cells, immune cells, etc. In some examples, each annotation includes an identifier associated with the annotator who generated the annotation.

[0151] In step 1304, the annotation processing system identifies a first annotation generated by a first annotator from the plurality of annotations. The first annotation may identify a first target type of biological structure (e.g., a tumor cell). In some examples, the first annotator is associated with a first weight. For example, a senior pathologist with many years of experience may be assigned a weight value that is greater than a weight value assigned to a junior pathologist with little or no experience.

[0152] In step 1306, the annotation processing system identifies a second annotation generated by a second annotator from the plurality of annotations. The second annotation identifies a second target type of biological structure (e.g., an immune cell), which is different from the first target type. Thus, the first and second annotations correspond to the same object or region but indicate biological structures of different target types. In some examples, the second annotator is associated with a second weighting, which may be the same or different from the first weighting assigned to the first annotator.

[0153] In step 1308, the annotation processing system determines that the first weight is greater than the second weight. Referring back to the example above, the first annotator may be a senior pathologist who is assigned a first weight that is greater than the second weight assigned to the second annotator, who may be a junior pathologist.

[0154] In step 1310, the annotation processing system generates a ground truth label for the first object or region of the image in response to determining that the first weight is greater than the second weight. The ground truth label may identify a first target type of biological structure. Continuing with the above example, the ground truth label may include a "tumor cell" label. The ground truth label at the consensus location may be used to train a machine learning model that predicts the type of biological structure in the corresponding region of another slide image. Training the machine learning model using the ground truth label with the consensus label and location may correspond to the training process described in FIG. 3. In some examples, the image with the generated ground truth label is displayed in a graphical user interface.

[0155] In some examples, the annotation processing system evaluates whether a ground truth label can be generated for the first object or region based on a confidence threshold. For example, the annotation processing system can determine a statistical value (e.g., average, sum) based on the first weight and the second weight and determine whether the statistical value exceeds the confidence threshold. In some examples, the confidence threshold includes a value that can be adjusted by one or more users. In response to determining that the statistical value does not exceed the confidence threshold, the annotation processing system can generate an output indicating that a consensus on the ground truth label was not reached between the first annotator and the second annotator. Thus, comparing the weights to the confidence threshold can filter out annotations with low weighted values, which may indicate that the consensus label is not sufficiently accurate.

[0156] In some examples, the annotation processing system utilizes the annotators' assigned weights to eliminate annotation-label conflicts that occur among three or more annotations. The annotation processing system identifies a third annotation generated by a third annotator from the multiple annotations. The third annotation identifies a second target type of biological structure (e.g., an immune cell), and the third annotator is associated with a third weight. The annotation processing system generates another statistical value based on the second weight and the third weight. Because both the second and third annotations identify the same target type of biological structure, another statistical value is generated based on the second weight and the third weight. In some examples, the other statistical value corresponds to the sum of the second and third weights. The annotation processing system determines that the other statistical value is greater than the first weight, and the annotation processing system can redefine the ground truth label of the object or region of the first image to identify the second target type of biological structure. Process 1300 then ends.

[0157] VII. Further Considerations Some embodiments of the present disclosure include a system including one or more data processors. In some embodiments, the system includes a non-transitory computer-readable storage medium including instructions that, when executed on the one or more data processors, cause the one or more data processors to perform some or all of one or more methods and / or some or all of one or more processes disclosed herein. Some embodiments of the present disclosure include a computer program product tangibly embodied in a non-transitory machine-readable storage medium, the computer program product including instructions configured to cause one or more data processors to perform some or all of one or more methods and / or some or all of one or more processes disclosed herein.

[0158] The terms and expressions which have been employed are used as terms of description rather than of limitation, and there is no intention in the use of such terms and expressions to exclude equivalents of the features shown and described or portions thereof, but it is recognized that various modifications are possible within the scope of the invention as claimed. Thus, although the claimed invention has been specifically disclosed by embodiments and optional features, it will be understood that modifications and variations of the concepts disclosed herein may be employed by those skilled in the art, and that such modifications and variations are deemed to be within the scope of the invention as defined by the appended claims.

[0159] The following description provides only preferred exemplary embodiments and is not intended to limit the scope, applicability, or configuration of the present disclosure. Rather, the following description of preferred exemplary embodiments will provide those skilled in the art with an enabling description for implementing various embodiments. It will be understood that various changes may be made in the function and arrangement of elements without departing from the spirit and scope as set forth in the appended claims.

[0160] In the following description, specific details are given to provide a comprehensive understanding of the embodiments. However, it will be understood that the embodiments may be practiced without these specific details. For example, circuits, systems, networks, processes, and other components may be shown as components in block diagram form in order to avoid obscuring the embodiments in unnecessary detail. In other instances, well-known circuits, processes, algorithms, structures, and techniques may be shown without unnecessary detail in order to avoid obscuring the embodiments.

Claims

1. 1. A method comprising: accessing a plurality of annotations associated with an image at a position that identifies an object or region within the image, the plurality of annotations being used to generate ground truth labels for supervised learning, validation, and / or testing of a machine learning model, each annotation of the plurality of annotations being located within the training image at a position that identifies an object or region within the training image, and each annotation including an identifier associated with an annotator that generated the annotation; identifying, from the plurality of annotations, a first annotation generated by a first annotator and a second annotation generated by a second annotator; determining a first distance value between a first position of the first annotation and a second position of the second annotation; determining that the first distance value is below a predetermined threshold; determining that the first distance value is below the predetermined threshold, wherein the first annotation and the second annotation both identify a first object or region within the image; identifying a third annotation from the plurality of annotations generated by the first annotator, the third annotation being located at a third location different from the first location; determining that a second distance value between the second position of the second annotation and the third position of the third annotation is below the predetermined threshold; In response to determining that the second distance value is below the predetermined threshold, also determining the third annotation as identifying the first object or region within the image; and in response to determining the third annotation as identifying the first object or region in the image, generating an output indicating that the object or region in the image has an annotation conflict; A method comprising:

2. determining that the first distance value is less than the second distance value; determining the third annotation as identifying a second object or region within the image; and generating another output in which the annotation conflicts are resolved; The method of claim 1 further comprising:

3. identifying a fourth annotation from the plurality of annotations generated by the second annotator, the fourth annotation being located at a fourth position on the image that is different from the second position of the second annotation; determining a third distance value between the third position and the fourth position; determining that the third distance value is below the predetermined threshold; determining that both the third annotation and the fourth annotation identify the first object or region within the image in response to determining that the third distance value is below the predetermined threshold; applying a clustering algorithm to the first location, the second location, the third location, and the fourth location to determine that (i) the first annotation and the second annotation correspond to a first sub-region of the first object or region, and (ii) the third annotation and the fourth annotation correspond to a second sub-region of the first object or region; 3. The method of claim 1 or claim 2, further comprising:

4. The method of claim 3 , wherein the clustering algorithm comprises a k-means clustering algorithm.

5. The method of any one of claims 1 to 4, wherein the first object or region depicts a biological structure.

6. the first annotation further identifies a target type of the biological structure depicted in the first object or region; the second annotation further identifies the target type of the biological structure depicted in the first object or region. The method of claim 5.

7. generating a ground truth label for the first object or region in the image, the ground truth label including the target type of the biological structure; and assigning the ground truth labels to the images and performing supervised learning of the machine learning model to predict whether another biological structure depicted in another image corresponds to the target type of the biological structure; The method of claim 6 further comprising:

8. 7. The method of claim 6, wherein the target type of biological structure corresponds to a stained tumor cell, an unstained tumor cell, or a normal cell.

9. 1. A method comprising: receiving a plurality of annotations corresponding to a first object or region of an image, the image depicting at least a portion of a tissue sample, the plurality of annotations being used to generate ground truth labels for supervised learning, validation, and / or testing of a machine learning model, each annotation in the set of annotations identifying a target type of a biological structure depicted in the first object or region, and each annotation including an identifier associated with an annotator that generated the annotation; identifying a first annotation from the plurality of annotations generated by a first annotator, the first annotation identifying a first target type of the biological structure, the first annotation being associated with a first weight; identifying a second annotation from the plurality of annotations, the second annotation generated by a second annotator, the second annotation identifying a second target type of the biological structure, the second annotator being associated with a second weight, and the first target type being different from the second target type; determining that the first weight is greater than the second weight; generating a ground truth label for the first object or region of the image in response to determining that the first weight is greater than the second weight, the ground truth label identifying the first target type of the biological structure; A method comprising:

10. determining a statistical value based on the first weight and the second weight; determining whether the statistical value exceeds a confidence threshold; 10. The method of claim 9, further comprising:

11. 11. The method of claim 10, further comprising: in response to determining that the statistical value does not exceed the confidence threshold, generating an output indicating that a consensus on the ground truth labels was not reached between the first annotator and the second annotator.

12. determining that the statistical value exceeds a confidence threshold; and in response to determining that the statistical value exceeds the confidence threshold, generating an output indicating that a consensus on the ground truth label has been reached between the first annotator and the second annotator; The method of claim 10 further comprising:

13. receiving input from a user; and adjusting the first weight and / or the second weight based on the input; The method according to any one of claims 9 to 12, wherein

14. identifying a third annotation from the plurality of annotations, the third annotation generated by a third annotator, the third annotation identifying the second target type of the biological structure, the third annotation being associated with a third weight; generating another statistical value based on the second weight and the third weight; determining that the other statistical value is greater than the first weight; and redefining the ground truth label of the first object or region of the image as identifying the second target type of the biological structure in response to determining that the other statistical value is greater than the first weight; The method of any one of claims 9 to 13, further comprising:

15. 15. The method of claim 9, further comprising: assigning the ground truth labels to the images for training the machine learning model, wherein the machine learning model is trained to predict whether another biological structure depicted in another image corresponds to the first target type of biological structure.

16. The method of any one of claims 9 to 15, further comprising displaying the images with the ground truth labels in a graphical user interface.

17. 1. A system comprising: one or more data processors; a non-transitory computer readable storage medium containing instructions that, when executed on said one or more data processors, cause said one or more data processors to perform any of the steps of the method of any of claims 1 to 16; and A system comprising:

18. A computer program product tangibly embodied in a non-transitory machine-readable storage medium comprising instructions configured to cause one or more data processors to perform any of the steps of the method of any one of claims 1 to 16.