Consensus markers in digital pathology images

By determining the consensus location and label of the annotation set in digital pathological images, the problem of pathologists' annotation inconsistency is solved, and the accuracy and consistency of machine learning models are improved.

CN120283269APending Publication Date: 2025-07-08VENTANA MEDICAL SYSTEMS INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202380082086.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-12-02
Filing Date
2023-11-27
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

In the prior art, machine learning models have inconsistent labeling of cell types in digital pathology, resulting in conflicting real labels in the training data, affecting the classification accuracy of the model.

Method used

Through the annotation processing system, the consensus location and consensus tags of the annotation sets at different locations of the image are determined, the annotation-position and annotation-label conflicts are resolved, and the true tags of consistency and accuracy are generated for training the model.

Benefits of technology

Improves the classification accuracy and consistency of machine learning models in digital pathological images, and reduces the impact of pathologist annotation errors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure HDA0005423611520000011
    Figure HDA0005423611520000011
  • Figure HDA0005423611520000021
    Figure HDA0005423611520000021
  • Figure HDA0005423611520000031
    Figure HDA0005423611520000031
Patent Text Reader

Abstract

The present disclosure relates to systems and methods for determining consensus locations and tags for a set of annotations associated with an object or region within an image. An annotation processing system may access a plurality of annotations associated with the image depicting at least a portion of the biological sample. The annotation processing system may determine a consensus position of a set of annotations positioned in different positions within a region of the image. At the determined consensus position, a consensus tag may be determined for the set of annotations identifying different target types of a biological structure. The consensus tags across different locations may be used to generate real tags for the image. The real tags may be used to train a machine learning model configured to predict different types of biological structures in a digital pathological image.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Priority Claim

[0002] This application claims priority to U.S. Provisional Patent Application No. 63 / 385,837, filed on Dec. 2, 2022, and entitled "CONSENSUS LABELING IN DIGITAL PATHOLOGY IMAGES", which is incorporated herein by reference in its entirety. Technical Field

[0003] The present disclosure generally relates to generating ground truth labels for images used to train, validate, or test a machine learning model. More specifically but without limitation, the present disclosure relates to generating consensus ground truth labels for images. Background Art

[0004] In a digital pathology platform, a machine learning model can be used to predict the cell type of various cells depicted in a slide image. The target cell type of a cell can include tumor cells, immune cells, stromal cells, etc. The machine learning model can be trained to predict the cell type based on a training data set including a plurality of training images. Each training image can include a set of ground truth labels, where the ground truth labels identify the target cell type of the cells at the corresponding positions in the training image. The set of ground truth labels can include one or more target classifications such that the machine learning model learns to predict one of the target classifications in the various cells detected in the slide image. The set of ground truth labels is typically provided by one or more experts (e.g., pathologists).

[0005] However, providing accurate ground truth labels can be challenging. For example, a first pathologist may view a particular cell as corresponding to a first cell type (e.g., an immune cell), but a second pathologist may view the same particular cell as corresponding to a second cell type (e.g., a tumor cell). These different observations can propagate to all other cells depicted in the training image, resulting in training data including conflicting ground truth labels. Such training data can include additional annotation conflicts because additional experts can be used to provide ground truth labels across various training images. If such training data is used for training, the machine learning model may not learn correctly and thus perform inconsistently during the inference phase (i.e., during deployment). Therefore, the classification accuracy of the trained machine learning model may be reduced. Summary of the Invention

[0006] The techniques described herein relate to determining a consensus position and label for a set of annotations associated with an object or region within an image. An annotation processing system can access a plurality of annotations associated with an image depicting at least a portion of a biological sample. The annotation processing system can determine a consensus position of the set of annotations at different positions located within a region of the image. At the determined consensus position, a consensus label can be determined for groups of annotations of different target types that identify biological structures. The consensus labels across different positions can be used to generate a ground truth label for the image.

[0007] In some embodiments, a system is provided that includes one or more data processors and a non-transitory computer-readable storage medium containing instructions that, when executed on the one or more data processors, cause the one or more data processors to perform part or all of one or more of the methods disclosed herein.

[0008] In some embodiments, a computer program product is provided that is tangibly embodied in a non-transitory machine-readable storage medium and includes instructions configured to cause one or more data processors to perform part or all of one or more of the methods disclosed herein.

[0009] The terms and expressions employed are used in a descriptive rather than a restrictive sense, and in using such terms and expressions, no intention is made to exclude any equivalents of the features shown and described or portions thereof, but it is recognized that various modifications are possible within the scope of the invention as claimed. Accordingly, it should be understood that although the invention as claimed has been specifically disclosed by way of embodiments and optional features, modifications and variations of the concepts disclosed herein may be employed by those skilled in the art, and such modifications and variations are considered to be within the scope of the invention as defined by the appended claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] This patent or application file contains at least one color drawing. Copies of this patent or patent application publication with one or more color drawings will be provided by the Patent Office upon request and payment of the necessary fee.

[0011] Aspects and features of various embodiments will become more apparent by reference to the drawings, in which:

[0012] Figure 1 An exemplary schematic diagram is shown that illustrates a process for generating a consensus label for a slide image according to some embodiments.

[0013] Figure 2 An exemplary network for generating a digital pathology image is shown according to various embodiments.

[0014] Figure 3Shows an exemplary computing environment for processing digital pathology images using machine learning / deep learning models according to various embodiments.

[0015] Figure 4 Shows a schematic diagram of a process for identifying clusters of sets of annotations for a slide image according to some embodiments.

[0016] Figure 5 Shows an exemplary image including the location of annotations according to some embodiments.

[0017] Figure 6 Shows an exemplary set of annotations identified in a training slide image according to some embodiments.

[0018] Figure 7 Shows an exemplary image of a process for resolving annotation conflicts according to some embodiments.

[0019] Figure 8 Shows another set of exemplary images of a process for removing annotation conflicts according to some embodiments.

[0020] Figure 9 Shows a process for identifying annotation-location conflicts in an image according to some embodiments.

[0021] Figure 10 Shows a process for removing annotation-location conflicts and determining a consensus label in an image according to some embodiments.

[0022] Figure 11 Shows a schematic diagram describing a process for removing annotation-label conflicts and determining a consensus label for an image according to some embodiments.

[0023] Figure 12 Shows an exemplary image including a consensus label and location according to some embodiments.

[0024] Figure 13 Shows a schematic diagram describing a process for removing annotation-label conflicts and determining a consensus label for an image according to some embodiments. Detailed Description

[0025] The prior art can generally remove annotation conflicts in the following ways: (i) receiving annotations for a given image region from multiple pathologists; and (ii) determining a single label for the image region based on the majority of annotations of a specific type indicating a biological structure. However, when it is unclear whether one or more annotations correspond to the same image region or image object (e.g., a cell), the above prior art may become ineffective. For example, although pathologists are instructed to annotate a cell by placing a dot at the center of the cell, the annotations can be inadvertently added (i.e., placed) at various positions that do not correspond to the center of the cell. In another example, some annotations may not be added to the image region at all because one or more pathologists may consider the image region as not depicting a cell. Therefore, due to inconsistent addition or non-addition of annotations at different positions in the image, it is challenging to determine an accurate ground truth label for the image based on processing the annotations.

[0026] Certain embodiments described herein can address these and other problems by determining a consensus position of a set of annotations located at different positions within a region of an image. The image can be part of a training data set for training a machine learning model. In some cases, the image is part of a validation data set for validating a machine learning model. Additionally or alternatively, the image can be part of a test data set for testing a machine learning model. For the determined consensus position, a consensus label can be determined for the set of annotations identifying different target types and / or positions of biological structures. The consensus label across different positions can be used to generate a ground truth label for the image. The ground truth label can be used to train a machine learning model configured to predict different types and / or positions of biological structures in digital pathology images. In fact, determining the consensus position and label can increase the consistency and accuracy of generating ground truth labels for training a machine learning model.

[0027] An annotation processing system can access multiple annotations associated with an image depicting at least a portion of a biological sample. As an illustrative example, the image can be a digital pathology image depicting at least a portion of a lung tissue sample obtained from a subject. The multiple annotations can be used to generate a ground truth label for supervised training of a machine learning model, where each annotation can include: (i) a location identifying an object or region within the image; (ii) a target type of the biological structure of the object or region; and (iii) an identifier associated with the annotator (e.g., a pathologist) who generated the annotation. Continuing with the above example, the annotation processing system can receive annotations of an image generated by a team of pathologists (e.g., four pathologists). The annotations in this example can include a set of x-y coordinates identifying the location of the image region, stained tumor cells (e.g., TC+) as the target type of the biological structure, and "Pathologist 4" as the identifier of the annotator who generated the annotation.

[0028] The annotation processing system can determine an annotation - location conflict for one or more annotations associated with a region or object of an image. In some cases, the annotation - location conflict is determined based on a region having two or more annotations from a single annotator. Continuing from the above example, the annotation processing system can identify an annotation - location conflict for a particular region by determining that five annotations are associated with a particular image region and further determining that two of the five annotations are submitted by an annotator identified as "Pathologist 4".

[0029] Then, the annotation processing system can resolve the annotation - location conflict for the region or object of the image. For each of the two or more annotations for a region associated with the same annotator, the annotation processing system can determine a statistical value (e.g., mean, median) representing the distance from the annotation to other annotations within the corresponding cluster. The statistical values between the two or more annotations from the same annotator can be compared. The annotation with the lowest statistical value can be determined to be part of the region, and the remaining conflicting annotations can be removed from being part of the region. Additionally or alternatively, the annotation processing system can apply a clustering algorithm (e.g., k - means algorithm) to group the annotations such that a first subset of annotations represents a first sub - region of the region and a second subset of annotations represents a second sub - region of the region of the image.

[0030] Continuing the above example, the annotation processing system can determine a first statistical value of 10 pixels for a first annotation generated by "Pathologist 4", where the first statistical value corresponds to the mean distance between the first annotation and other annotations present in the particular region. The annotation processing system can also determine a second statistical value of 15 pixels for a second annotation generated by "Pathologist 4", where the second statistical value corresponds to the mean distance between the second annotation and other annotations. The annotation processing system can resolve the annotation - location conflict by discarding the second annotation with the larger statistical value.

[0031] After determining the consensus position for an annotation set, the annotation processing system can determine the ground truth label for the consensus position of the image. In some cases, if all the annotations in the annotation set indicate the same target type of biological structure (e.g., tumor cells), the ground truth label is determined. However, in some cases, there is no clear consensus among the annotation sets regarding the target type of the biological structure. In such cases, the annotation processing system can resolve the "annotation-label" conflict by: (i) determining the weights associated with each annotation; and (ii) comparing the weights to determine the consensus label for the consensus position. The weights are determined based on the performance and historical data associated with the annotators who generated the annotations. For example, the annotator of the first annotation can be a senior pathologist with several years of experience. The weight of the first annotation can be assigned a value relatively greater than another weight of a second annotation associated with a less experienced pathologist.

[0032] Continuing with this example, the consensus position includes four annotations, where the annotators "Pathologist 1" and "Pathologist 3" mark the position as having stained tumor cells, while the other two annotators "Pathologist 2" and "Pathologist 4" mark the position as having unstained tumor cells. Then, the first sum of the weights assigned to "Pathologist 1" and "Pathologist 3" can be compared with the second sum of the weights assigned to "Pathologist 2" and "Pathologist 4". If it is determined that the first sum is greater than the second sum, the annotation processing system can determine the consensus label of the consensus position as "stained tumor cells".

[0033] Then, the annotation processing system can generate a ground truth label set for the image based at least in part on the consensus position and the corresponding consensus label. Then, this ground truth label set of the image can be used to perform supervised training of a machine learning model. In some cases, this ground truth label set of the image is used to validate the machine learning model. Additionally or alternatively, this ground truth label set of the image can be used to test the machine learning model. By determining the consensus position and label, accurate and consistent performance of the machine learning model can be achieved.

[0034] In some cases, the annotation processing system optimizes the determination of the consensus label based on various adjustable parameters. For example, the annotation processing system can adjust the weights based on the performance of the corresponding annotators across different images or multiple FOVs of a single image. In addition to adjusting the weights, other types of customizations can be applied by the user to determine the consensus position and label. For example, a distance threshold can be adjusted by the user to determine whether the annotation sets identify similar positions in the image. These customizations allow for a better representation of the data for calculating the consensus. Further, the annotations and any conflicts between the annotations can be displayed on a graphical user interface where the user can adjust the weight values and the distance threshold.

[0035] Certain embodiments described herein improve the training of machine learning models for classifying biological structures in digital pathology images. An annotation processing system can improve the consistency and accuracy of generating ground truth labels by facilitating resolution of various types of conflicts across the annotations in an image (e.g., annotation-location conflicts, annotation-label conflicts). In fact, the annotation processing system can be enabled to mitigate errors introduced by annotators (e.g., pathologists), such as making a particular annotated click at the wrong location in the image when generating an annotation. Additionally, using adjustable weights for annotators can fine-tune the consistency and accuracy of the ground truth labels generated for an image. Thus, the embodiments herein reflect improvements in the functionality of artificial intelligence systems and digital pathology image processing techniques.

[0036] Although certain embodiments are described, these embodiments are presented by way of example only and are not intended to limit the scope of protection. The devices, methods, and systems described herein can be embodied in many other forms. Additionally, various omissions, substitutions, and changes can be made to the forms of the example methods and systems described herein without departing from the scope of protection.

[0037] I. Definitions

[0038] As used herein, when an action is “based on” something, this means that the action is based at least in part on at least a portion of something.

[0039] As used herein, the terms “substantially,” “about,” and “approximate” are defined as being largely but not necessarily completely as specified as would be understood by a person of ordinary skill in the art (and including completely as specified). In any of the disclosed embodiments, the terms “substantially,” “about,” or “approximate” can be replaced with “within [a certain percentage]” for the specified, where the percentage includes 0.1%, 1%, 5%, and 10%.

[0040] As used herein, the terms "sample", "biological sample", "tissue", or "tissue sample" refer to any sample obtained from any living organism, including viruses, that includes biomolecules such as proteins, peptides, nucleic acids, lipids, carbohydrates, or combinations thereof. Examples of other living organisms include mammals (such as humans; veterinary animals such as cats, dogs, horses, cows, and pigs; and laboratory animals such as mice, rats, and primates), insects, annelids, arachnids, marsupials, reptiles, amphibians, bacteria, and fungi. Biological samples include tissue samples (such as tissue sections and needle biopsies of tissue), cell samples (such as cytology smears such as cervical smears or blood smears or cell samples obtained by microdissection) or cell fractions, fragments, or organelles (such as obtained by lysing cells and separating their components by centrifugation or otherwise). Other examples of biological samples include blood, serum, urine, semen, feces, cerebrospinal fluid, interstitial fluid, mucus, tears, sweat, pus, biopsy tissue (e.g., obtained by surgical biopsy or needle biopsy), nipple aspirate, cerumen, milk, vaginal secretions, saliva, swabs (such as oral swabs), or any material containing biomolecules and derived from a first biological sample. In some embodiments, the term "biological sample" as used herein refers to a sample prepared from a tumor or a part thereof obtained from a subject (such as a homogenized or liquefied sample).

[0041] As used herein, the terms "biomaterial", "biological structure", or "cellular structure" refer to natural materials or structures that contain whole or part of a living structure (e.g., cell nucleus, cell membrane, cytoplasm, chromosome, DNA, cell, cell cluster, etc.).

[0042] As used herein, a "digital pathology image" refers to a digital image of a stained sample.

[0043] As used herein, the term "cell detection" refers to the detection of the pixel locations and features of cells or cellular structures (e.g., cell nucleus, cell membrane, cytoplasm, chromosome, DNA, cell, cell population, etc.).

[0044] As used herein, the terms "region" or "object" refer to a set of pixels of an image that includes image data and that is intended to be evaluated during an image analysis process. A region or object can depict a tissue region (e.g., tumor cells or staining expression) that is intended to be analyzed during an image analysis process.

[0045] As used herein, the term "field of view" or "FOV" refers to a whole slide image or a portion of a whole slide. In some embodiments, the "FOV" refers to the area scanned of a whole slide or a target area having (x, y) pixel dimensions (e.g., 1000 pixels x 1000 pixels). In some embodiments, the FOV includes the same pixel resolution and / or magnification level associated with the corresponding whole slide image. Additionally or alternatively, the FOV includes different pixel resolutions and / or magnification levels associated with the corresponding whole slide image.

[0046] II. Overview

[0047] Figure 1 Exemplary schematic diagram 100 showing a process for generating consensus labels for slide images according to some embodiments is shown. Schematic diagram 100 includes: (i) an annotation processing system 102 configured to generate consensus positions and labels for a given image; and (ii) a graphical user interface 116 configured to display intermediate outputs of the annotation processing system 102 to a user.

[0048] In block 104, the annotation processing system 102 receives a set of annotations associated with an image. The image can be a digital pathology image depicting at least a portion of a biological sample (e.g., tissue) obtained from a subject. In some cases, an annotation tool is used by an annotator to provide the set of annotations to the whole tissue image or a selected region of the image (e.g., a field of view (FOV)). The annotation tool can aggregate the set of annotations into a schema file (e.g., an XML document) and transmit the schema file to the annotation processing system 102 for further processing. In some cases, the annotation tool is provided to each of a plurality of annotators involved in generating ground truth labels for the image, where the corresponding schema files are transmitted to the annotation processing system 102 for further processing. For example, the annotation processing system 102 can receive 5 schema files, each of which includes a set of annotations submitted by a corresponding annotator (e.g., a pathologist). In some embodiments, the annotation processing system 102 receives and processes more than 10, 15, 20, 25, 30, 35, 40, 45, 50, or more than 50 schema files for the corresponding image.

[0049] The annotation tool can also provide available tags for classifying cells in a given image. An annotation can be added to a location in the image, where the annotation can include a tag identifying the type of cell depicted in the image. In some cases, the tags include: (i) a TC+ tag identifying tumor cells stained with a specific biomarker; (ii) a TC− tag identifying another tumor cell that is not stained; (iii) an IC tag identifying immune cells; and (iv) another tag identifying cells other than tumor cells or immune cells.

[0050] In response to receiving the set of annotations, the annotation processing system 102 can parse each annotation in the set to identify its corresponding annotation data. For example, an annotation can be parsed by the annotation processing system 102 to identify an initial location of a region of the image and an initial tag (e.g., a stained tumor cell) classifying a biological structure depicted at the initial location of the image. In addition to the initial location and tag, the annotation processing system 102 can identify an identifier of the annotator who submitted the annotation, the type of tissue depicted by the image, the number of annotators who submitted the set of annotations for the image, and the tags used to classify the cells of the image.

[0051] In block 106, the annotation processing system 102 identifies one or more annotation clusters associated with the image. For example, the location of an annotation can include a set of coordinates of the image that the annotator has indicated as depicting a specific biological structure. The cluster can be defined as a set of annotations within a predetermined distance. In some cases, a threshold is used to determine the set of annotations, where the threshold identifies the number of pixels between annotations (e.g., 10 pixels). The threshold can be adjusted by the annotator, which can increase or decrease the number of annotations to be included in the cluster. By adjusting the threshold to increase or decrease the number of annotations, the size of the cluster in the image can be defined.

[0052] In block 108, the annotation processing system 102 identifies annotation-location conflicts of the annotations. The annotation processing system 102 can identify an annotation-location conflict by determining that two or more annotations associated with the same annotator (e.g., a pathologist) are associated within a specific cluster. Additionally or alternatively, the annotation processing system 102 can identify an annotation-location conflict by detecting that the number of annotations in a specific cluster (e.g., 6 annotations) exceeds the total number of annotators who submitted the annotations for the image (e.g., 5 pathologists). In some cases, the graphical user interface 116 displays the annotations having the annotation-location conflict (block 118).

[0053] In block 110, the annotation processing system 102 resolves annotation-location conflicts. For annotations that include location conflicts, the annotation processing system 102 can determine a statistical value (e.g., mean, median) representing the distance from the annotation to other annotations within the corresponding region for each of the two or more annotations having the conflict (e.g., annotations generated by a particular annotator for the same object). The statistical values between the two or more annotations from the same annotator can be compared. The annotation with the lowest statistical value can be determined to be part of the annotation for the region, while the remaining conflicting annotations can be removed from the association with the region. In some cases, the graphical user interface 116 displays the annotations for which the annotation-location conflicts have been resolved (block 120).

[0054] In block 112, the annotation processing system 102 determines the consensus location of one or more regions of the image. Each region can include a set of annotations submitted by multiple annotators, and the consensus location of the cluster can be determined based on the median coordinates among the locations of the set of annotations associated with the cluster. In some cases, the count of the annotations for the region is compared with a predetermined confidence threshold. If the count is below the confidence threshold, the set of annotations is discarded from the image. Thus, no ground truth label is generated for such regions.

[0055] In block 114, the annotation processing system 102 determines the consensus label for the image. Even when there is no majority label for a given consensus location, the annotation processing system 102 can use weighted values to determine the consensus label. For example, the consensus location can include four annotations, where two annotations include labels identifying the location as depicting tumor cells and the other two annotations include other labels identifying the location as depicting normal cells. Based on the weights applied to each of the annotations, the annotation processing system 102 can resolve the "ties" among the annotations and determine the consensus label for the location. In some cases, the graphical user interface 116 displays the consensus location and label for the image (block 122).

[0056] In some cases, the graphical user interface 116 allows one or more parameters used by the annotation processing system 102 to be adjusted in response to one or more of visualizations 118, 120, and 122. For example, the graphical user interface 116 can adjust the weighted value assigned to an annotator associated with a particular set of annotations. The weighted value can be adjusted by determining a measure of the annotator's accuracy based on the validation of an image with a ground truth label. The measure of accuracy can be determined based on the percentage of annotations submitted by the annotator that correspond to the ground truth (e.g., 75%, 85%). In some cases, the weighted value of the annotator is adjusted in proportion to the measure of accuracy determined for the annotator.

[0057] III. Generating Digital Pathology Images

[0058] Digital pathology involves the interpretation of digitized images to correctly diagnose a subject and guide treatment decisions. In digital pathology solutions, an image analysis workflow can be established to automatically detect or classify biological objects of interest, such as positive and negative tumor cells, etc. Exemplary digital pathology solution workflows include obtaining a tissue slide, scanning a preselected area or the entire slide using a digital image scanner (e.g., a whole slide image (WSI) scanner) to obtain a digital image, performing image analysis on the digital image using one or more image analysis algorithms, and possibly detecting and quantifying each object of interest (e.g., counting or identifying the object-specific or cumulative area of each object of interest) based on the image analysis (e.g., quantitative or semi-quantitative scoring, such as positive, negative, moderate, weak, etc.).

[0059] Figure 2 An exemplary network 200 for generating digital pathology images is shown. A fixation / embedding system 205 uses a fixative (e.g., a liquid fixative such as a formaldehyde solution) and / or an embedding substance (e.g., histological wax such as paraffin and / or one or more resins such as styrene or polyethylene) to fix and / or embed a tissue sample (e.g., a sample including at least a portion of at least one tumor). Each sample can be fixed by exposing the sample to the fixative for a predetermined period (e.g., at least 3 hours) and then dehydrating the sample (e.g., via exposure to an ethanol solution and / or a clearing agent). When the sample is in a liquid state (e.g., when heated), the embedding substance can infiltrate the sample.

[0060] Sample fixation and / or embedding is used to preserve the sample and slow down sample degradation. In histology, fixation generally refers to an irreversible process of using chemicals to retain chemical components, preserve the natural sample structure, and keep the cell structure from degrading. Fixation may also harden the cells or tissue for sectioning. Fixatives can enhance the preservation of samples and cells using cross-linked proteins. Fixatives may bind to and cross-link some proteins and denature other proteins through dehydration, which may harden the tissue and inactivate enzymes that might otherwise degrade the sample. Fixatives can also kill bacteria.

[0061] The fixative can be applied, for example, by perfusion and infiltration of the prepared sample. A variety of fixatives can be used, including methanol, Bouin's fixative, and / or formaldehyde-based fixatives such as neutral buffered formalin (NBF) or paraffin-formalin (paraformaldehyde - PFA). In the case where the sample is a liquid sample (e.g., a blood sample), the sample can be smeared on a slide and dried before fixation. Although for histological research purposes, the fixation process can be used to preserve the structure of samples and cells, fixation may cause the hiding of tissue antigens, thereby reducing antigen detection. Therefore, fixation is generally considered a limiting factor in immunohistochemistry because formalin can crosslink antigens and mask epitopes. In some cases, additional processes are carried out to reverse the effects of crosslinking, including treating the fixed sample with citraconic anhydride (a reversible protein crosslinker) and heating.

[0062] Embedding can include infiltrating the sample (e.g., a fixed tissue sample) with a suitable histological wax such as paraffin. Histological wax may be insoluble in water or alcohol but soluble in paraffin solvents such as xylene. Therefore, the water in the tissue may need to be replaced with xylene. For this purpose, the tissue can be dehydrated first by gradually replacing the water in the sample with alcohol, which can be achieved by passing the tissue through increasing concentrations of ethanol (e.g., from 0% to about 100%). After replacing the water with alcohol, the alcohol can be replaced with xylene, which is miscible with alcohol. Since histological wax is soluble in xylene, the melted wax may fill the space filled with xylene and previously filled with water. The sample filled with wax can be cooled to form a hardened block, which can be clamped into a microtome, vibratome, or compression vibratome for sectioning. In some cases, deviating from the above example procedures may result in paraffin infiltration, thus inhibiting the penetration of antibodies, chemicals, or other fixatives.

[0063] Then, a microtome 210 can be used to section the fixed and / or embedded tissue sample (e.g., a tumor sample). Sectioning is the process of cutting thin slices (e.g., with a thickness of, for example, 2 - 5 μm) of the sample from the tissue block for the purpose of fixing it on a microscope slide for examination. A microtome, vibratome, or compression vibratome can be used for sectioning. In some cases, the tissue can be rapidly frozen in dry ice or isopentane and then cut with a cold knife in a cold cabinet (e.g., a cryostat). Other types of coolants can be used to freeze the tissue, such as liquid nitrogen. Sections for bright-field and fluorescence microscopy are typically about 2 μm to 10 μm thick. In some cases, the sections can be embedded in epoxy resin or acrylic resin so that thinner sections (e.g., <2 μm) can be cut. Then the sections can be mounted on one or more slides. A coverslip can be placed on top to protect the sample sections.

[0064] Because tissue sections and the cells therein are actually transparent, the preparation of the sections typically further includes staining (e.g., automated staining) the tissue sections to make the relevant structures more visible. In some cases, the staining is performed manually. In some cases, a staining system 215 is used to perform the staining semi-automatically or automatically. The staining process includes exposing sections of a tissue sample or a fixed liquid sample to one or more different stains (e.g., sequentially or simultaneously) to reveal different features of the tissue.

[0065] For example, staining can be used to label specific types of cells and / or mark specific types of nucleic acids and / or proteins to assist microscopy. The staining process typically involves adding a dye or stain to the sample to identify or quantify the presence of a specific compound, structure, molecule, or feature (e.g., subcellular feature). For example, staining can help identify or highlight specific biomarkers in a tissue section. In other instances, stains can be used to identify or highlight biological tissues (e.g., muscle fibers or connective tissue), cell populations (e.g., different blood cells), or organelles within individual cells.

[0066] One exemplary type of tissue staining is histochemical staining, which uses one or more chemical dyes (e.g., acidic dyes, basic dyes, chromogens) to stain the tissue structure. Histochemical staining can be used to indicate general aspects of tissue morphology and / or cytohistology (e.g., to distinguish the nucleus from the cytoplasm, indicate lipid droplets, etc.). An example of a histochemical stain is H&E. Other examples of histochemical stains include trichrome stains (e.g., Masson's trichrome), periodic acid-Schiff (PAS), silver stains, and iron stains. The molecular weight of histochemical staining reagents (e.g., dyes) is typically about 500 kilodaltons (kD) or less, although some histochemical staining reagents (e.g., Alcian blue, phosphomolybdic acid (PMA)) may have molecular weights as high as two or three thousand kD. An example of a high-molecular-weight histochemical staining reagent is α-amylase (about 55 kD), which can be used to indicate glycogen.

[0067] Another type of tissue staining is IHC (also known as "immunostaining"), which uses a primary antibody that specifically binds to a target antigen of interest (also known as a biomarker). IHC can be direct or indirect. In direct IHC, the primary antibody is directly conjugated to a label (e.g., a chromophore or a fluorophore). In indirect IHC, the primary antibody first binds to the target antigen, and then a secondary antibody conjugated to a label (e.g., a chromophore or a fluorophore) binds to the primary antibody. The molecular weight of IHC reagents is much higher than that of histochemical staining reagents because the molecular weight of an antibody is about 150 kD or higher.

[0068] A variety of staining protocols can be used for staining. For example, exemplary IHC staining protocols include: using a hydrophobic barrier line around the sample (e.g., tissue slide) to prevent reagent leakage from the slide during incubation; treating the tissue slide with reagents to block endogenous sources of non-specific staining (e.g., enzymes, free aldehyde groups, immunoglobulins, other unrelated molecules that can mimic specific staining); incubating the sample with a permeabilization buffer to facilitate the penetration of antibodies and other staining reagents into the tissue; incubating the tissue slide with a primary antibody at a specific temperature (e.g., room temperature, 6°C - 8°C) for a period of time (e.g., 1 hour to 24 hours); rinsing the sample with a wash buffer; then incubating the sample (tissue slide) with a secondary antibody at another specific temperature (e.g., room temperature) for another period of time; rinsing the sample again with a water buffer; incubating the rinsed sample with a chromogen (e.g., DAB: 3,3'-diaminobenzidine); and washing away the chromogen to stop the reaction. In some cases, counterstaining is subsequently used to identify the overall "landscape" of the sample and as a reference for the primary color used to detect tissue targets. Examples of counterstains can include hematoxylin (which stains from blue to purple), methylene blue (which stains blue), toluidine blue (which stains cell nuclei dark blue and polysaccharides from pink to red), nuclear fast red (also known as Kernchtrot dye, which stains red), and methyl green (which stains green); non-nuclear chromogenic stains, such as eosin (which stains pink), etc. Those of ordinary skill in the art will recognize that other immunohistochemical staining techniques can be implemented for staining.

[0069] In another example, a tissue section can be stained with an H&E staining protocol. The H&E staining protocol includes applying a hematoxylin stain or mordant mixed with a metal salt to the sample. The sample can then be rinsed in a weak acid solution to remove excess stain (differentiation), and then blued in a weakly alkaline water. After applying hematoxylin, the sample can be counterstained with eosin. It should be understood that other H&E staining techniques can be implemented.

[0070] In some embodiments, various types of stains can be used for staining, depending on the target features being targeted. For example, DAB can be used for various tissue sections in IHC staining, where DAB produces a brown color depicting the target features in the stained image. In another example, alkaline phosphatase (AP) can be used for skin tissue sections in IHC staining because the DAB color may be masked by melanin. Regarding primary staining techniques, applicable stains can include, for example, basophilic and eosinophilic stains, hematoxylin and eosin, silver nitrate, trichrome stains, etc. Acidic dyes can react with cationic or basic components in tissues or cells, such as proteins and other components in the cytoplasm. Basic dyes can react with anionic or acidic components in tissues or cells, such as nucleic acids. As mentioned above, one example of a staining system is H&E. Eosin may be a negatively charged pink acidic dye, and hematoxylin may be a purple or blue basic dye, which includes hematein and aluminum ions. Other examples of stains can include periodic acid-Schiff reaction (PAS) stains, Masson's trichrome stains, Alcian blue stains, Van Gieson stains, reticular fiber stains, etc. In some embodiments, different types of stains can be used in combination.

[0071] The sections can then be mounted on corresponding glass slides, and then the imaging system 220 can scan or image to generate the original digital pathology images 225a to 225n. A microscope (e.g., an electron microscope or an optical microscope) can be used to magnify the stained sample. For example, the resolution of an optical microscope may be less than 1 μm, such as approximately a few hundred nanometers. To observe finer details in the nano or sub-nano range, an electron microscope can be used. An imaging device (combined with or separate from the microscope) images the magnified biological sample to obtain image data, such as a multi-channel image (e.g., multi-channel fluorescence) having multiple (such as, for example, ten to sixteen) channels. The imaging device can include, but is not limited to, a camera (e.g., an analog camera, a digital camera, etc.), optics (e.g., one or more lenses, sensor focusing lens groups, microscope objectives, etc.), an imaging sensor (e.g., a charge-coupled device (CCD), a complementary metal-oxide semiconductor (CMOS) image sensor, etc.), photographic film, etc. In a digital embodiment, the imaging device can include multiple lenses that can cooperate to demonstrate an instant focusing function. The imaging sensor (e.g., a CCD sensor) can capture a digital image of the biological sample. In some embodiments, the imaging device is a bright-field imaging system, a multi-spectral imaging (MSI) system, or a fluorescence microscopy system. The imaging device can utilize invisible electromagnetic radiation (e.g., UV light) or other imaging techniques to capture the image. For example, the imaging device can include a microscope and a camera arranged to capture the image magnified by the microscope. The image data received by the analysis system can be the same as and / or can be derived from the original image data captured by the imaging device.

[0072] The image of the stained section can then be stored in a storage device 225 such as a server. The images can be stored locally, remotely, and / or in a cloud server. Each image can be stored in association with an identifier of the subject and a date (e.g., the date the sample was collected and / or the date the image was captured). The images can be further transmitted to another system (e.g., a system associated with a pathologist, an automated or semi-automated image analysis system, or a machine learning training and deployment system, as described in further detail herein).

[0073] It should be understood that modifications to the processes described with respect to network 200 are contemplated. For example, if the sample is a liquid sample, embedding and / or sectioning can be omitted from the process.

[0074] IV. Exemplary Systems for Digital Pathology Image Conversion

[0075] Figure 3 A block diagram showing a computing environment 300 for processing digital pathology images using a machine learning model is shown. As described further herein, processing digital pathology images can include using digital pathology images to train a machine learning algorithm and / or using a trained (or partially trained) version of a machine learning algorithm (i.e., a machine learning model) to convert some or all of the digital pathology images into one or more results.

[0076] As Figure 3 shown, the computing environment 300 includes several stages: an image storage stage 305, a preprocessing stage 310, a labeling stage 315, a data augmentation stage 317, a training stage 320, and a result generation stage 325.

[0077] The image storage period 305 includes one or more image data stores 330 (e.g., the storage device 230 described with respect to Figure 2 ), which are accessed (e.g., by the preprocessing period 310) to provide a set of digital images 335 from a preselected area of a biological sample slide or an entire biological sample slide (e.g., a tissue slide). Each digital image 335 stored in each image data store 330 and accessed during the image storage period 310 can include a digital pathology image generated according to some or all of the processes described with respect to Figure 2 the network 200 depicted. In some embodiments, each digital image 335 includes image data from one or more scanned slides. Each of the digital images 335 can correspond to image data from a single sample and / or image data on the day the underlying image data corresponding to the image was collected.

[0078] Image data can include an image, as well as any information related to color channels or color wavelength channels, and details about the imaging platform on which the image is generated. For example, tissue sections may need to be stained by applying a staining assay that contains one or more different biological tags associated with chromogenic stains or fluorophores for brightfield imaging or fluorescence imaging. The staining assay can use chromogenic stains for brightfield imaging, organic fluorophores, quantum dots, or a combination of organic fluorophores and quantum dots for fluorescence imaging, or any other combination of stains, biological tags, and the observation or imaging device. Exemplary biomarkers include biomarkers such as estrogen receptor (ER), human epidermal growth factor receptor 2 (HER2), human Ki-67 protein, progesterone receptor (PR), programmed cell death protein 1 (PD1), etc., where the tissue section is detectably labeled with a binding agent (e.g., an antibody) for each of ER, HER2, Ki-67, PR, PD1, etc. In some embodiments, digital image and data analysis operations (such as classification, scoring, cox modeling, and risk stratification) depend on the type of biomarker used and the field of view (FOV) selection and annotation. Additionally, typical tissue sections are processed in an automated staining / platform that applies the staining assay to the tissue section, resulting in a stained sample. There are a variety of commercial products available on the market that are suitable for use as staining / assay platforms, an example being the product of assignee Ventana Medical Systems, Inc. The stained tissue section can be provided to an imaging system, such as a microscope or a whole slide scanner having a microscope and / or imaging components, an example being the product of assignee Ventana Medical Systems, Inc. iScan DP200. Multiplex tissue slides can be scanned on an equivalent multiplex slide scanner system. Additional information provided by the imaging system can include any information related to the staining platform, including the concentration of the chemicals used for staining, the reaction time of the chemicals applied to the tissue in the staining, and / or the pre-analytical conditions of the tissue, such as tissue age, fixation method, duration, how the section is embedded, cut, etc.

[0079] During preprocessing period 310, each of one, more, or all digital images in digital image set 335 is preprocessed using one or more techniques to generate corresponding preprocessed images 340. Preprocessing can include cropping the images. In some cases, preprocessing can further include normalizing or resizing (e.g., normalizing) to place all features on the same scale (e.g., the same size scale or the same color scale or color saturation scale). In some cases, the size of the image is adjusted using a minimum size (width or height) of a predetermined number of pixels (e.g., 2500 pixels) or a maximum size (width or height) of a predetermined number of pixels (e.g., 3000 pixels), and the original aspect ratio is optionally maintained. Preprocessing can further include removing noise. For example, the image can be smoothed, such as by applying a Gaussian function or Gaussian blur, to remove unwanted noise.

[0080] The preprocessed images 340 can include one or more training images, validation images, test images, and unlabeled images. It should be understood that it is not necessary to access the preprocessed images 340 corresponding to the training set, validation set, and unlabeled set simultaneously. For example, an initial set 340 of training and validation preprocessed images can be accessed first and used to train machine learning algorithm 355, and the unlabeled input images can subsequently be accessed or received (e.g., at a single or multiple subsequent times) and used by the trained machine learning model 360 to provide a desired output (e.g., cell classification).

[0081] In some cases, supervised training is used to train the machine learning algorithm 355, and during the labeling period 315, some or all of the pre-processed images 340 are labeled, at least in part, manually, semi-automatically, or automatically using the labels 345, which identify the "correct" interpretation (i.e., "ground truth") of the various biological materials and structures within the pre-processed images 340. For example, the label 345 can identify target features (e.g.), the classification of cells, a binary indication as to whether a given cell is a particular type of cell, a binary indication as to whether the pre-processed image 340 (or a particular region of the pre-processed image 340) includes a particular type of depiction (e.g., necrosis or artifact), a classification characterization of a slide-level or region-specific depiction (e.g., identifying a particular type of cell), a quantity (e.g., identifying the number of a particular type of cell within a region, the number of depicted artifacts, or the number of necrotic regions), the presence or absence of one or more biomarkers, etc. In some cases, the label 345 includes a location. For example, the label 345 can identify the point location of the nucleus of a particular type of cell or the point location of a particular type of cell (e.g., a raw point label). As another example, the label 345 can include a boundary or border, such as the boundary of a depicted tumor, blood vessel, necrotic region, etc. As another example, the label 345 can include one or more biomarkers identified based on a biomarker pattern observed using one or more stains. For example, a tissue slide stained for a biomarker such as programmed cell death protein 1 ("PD1") can be observed and / or processed such that cells are labeled as positive or negative cells based on the expression level and pattern of PD1 within the tissue. Depending on the target feature, a given pre-processed image 340 can be associated with a single label 345 or multiple labels 345. In the latter case, each label 345 can be associated with an indication (e.g.) as to which location or portion of the pre-processed image 345 the label corresponds to.

[0082] The label 345 assigned at the labeling stage 315 can be identified based on input from a human user (e.g., a pathologist or an image scientist) and / or an algorithm (e.g., an annotation tool) configured to define the label 345. In some cases, the labeling stage 315 can include transmitting and / or presenting some or all of the one or more preprocessed images 340 to a computing device operated by a user. In some cases, the labeling stage 315 includes being presented by a labeling controller 350 at a computing device operated by a user using an interface (e.g., using an API), where the interface includes input components to accept input identifying the label 345 for a target feature. For example, a user interface can be provided by the labeling controller 350 that enables selection of an image or an image region (e.g., FOV) for labeling. A user operating an operation terminal can use the user interface to select an image or an FOV. Several image or FOV selection mechanisms can be provided, such as specifying a known or irregular shape, or defining a target anatomical region (e.g., a tumor region). In one example, the image or FOV is a whole tumor region selected on an IHC slide stained with a combination of H&E stain. The image or FOV selection can be performed by a user or through an automated image analysis algorithm, such as segmentation of a tumor region on an H&E tissue slide. For example, a user can select an image or an FOV as a whole slide or a whole tumor, or can use a segmentation algorithm to automatically designate the whole slide or whole tumor region as an image or an FOV. Thereafter, a user operating an operation terminal can select one or more labels 345 to apply to the selected image or FOV, such as a point location on a cell, a positive marker for a biomarker expressed by the cell, a negative biomarker for a biomarker not expressed by the cell, a boundary around the cell, etc.

[0083] In some cases, the interface can identify which specific markers 345 are being requested and / or the degree to which specific markers are being requested, which can be communicated to the user via (e.g.) text instructions and / or visualization. For example, a specific color, size, and / or symbol can indicate that a label 345 is being requested for a specific depiction (e.g., a specific cell or region or staining pattern) within an image relative to other depictions. If labels 345 corresponding to multiple depictions are to be requested, the interface can identify each of the depictions simultaneously or can identify each depiction in sequence (such that providing a label for one identified depiction triggers the identification of the next depiction for labeling). In some cases, each image is presented until the user has identified a specific number of labels 345 (e.g., a specific type of label). For example, a given whole-slide image or a given patch of a whole-slide image can be presented until the user has identified the presence or absence of three different biomarkers, at which point the interface can present a different whole-slide image or an image of a different patch (e.g., until a threshold number of images or patches have been labeled). Thus, in some cases, the interface is configured to request and / or accept labels 345 for an incomplete subgroup of target features, and the user can determine which of the many possible depictions will be labeled.

[0084] In some cases, the labeling period 315 includes a labeling controller 350 that implements an annotation algorithm to semi-automatically or automatically label various features of an image or a target region within an image. The labeling controller 350 annotates the image or FOV on the first slide based on input from the user or the annotation algorithm and maps that annotation across the remainder of the slide. Depending on the defined FOV, a variety of methods for annotation and registration are possible. For example, it can be automated or done by the user in something such as VIRTUOSO / VERSO TMSelect the annotated whole tumor region on the H&E slide among multiple consecutive slides on the interface of the like. Since the other tissue slides correspond to consecutive sections from the same tissue block, the marker controller 350 performs an inter-marker registration operation to map and transfer the whole tumor annotation from the H&E slide to each of the remaining IHC slides in the series. Exemplary methods for inter-marker registration are further described in detail in the commonly assigned International Application WO2014140070A2, "Whole Slide Image Registration and Cross-Image Annotation Devices, Systems, and Methods," filed on March 12, 2014, which is hereby incorporated by reference in its entirety for all purposes. In some embodiments, any other method for image registration and generating whole tumor annotations may be used. For example, a qualified reader such as a pathologist may annotate the whole tumor region on any other IHC slide and perform the marker controller 350 to map the whole tumor annotation to other digital slides. For example, a pathologist (or an automated detection algorithm) may annotate the whole tumor region on the H&E slide, thereby triggering the analysis of all adjacent consecutive section IHC slides to determine the whole slide tumor score for the annotated regions on all slides.

[0085] In some cases, the annotation phase 315 further includes an annotation processing system 351 that implements an annotation algorithm to identify annotation-location and annotation-label conflicts within an annotation set associated with an image (or FOV of an image). The annotation processing system 351 may determine a consensus location for groups of annotations located at different positions within a training image region. In some cases, the annotation processing system 351 determines that an annotation-location conflict exists in a region by determining that two or more annotations from the same annotator exist in the region of the training image. The annotation processing system 351 may resolve such location conflicts by retaining the annotation closest to other annotations in the region while discarding other annotations from the same annotator. At the determined consensus location, a consensus label may be determined for groups of annotations identifying different target types of biological structures. The consensus labels across different positions may be used to generate a ground truth label for the image. The ground truth label may be used for training, validating, and / or testing a machine learning model configured to predict different types of biological structures in digital pathology images.

[0086] During enhancement period 317, the training images (original images), whether labeled or unlabeled, from the preprocessed image 340 are enhanced using the synthetic image 352 generated by the enhancement control 354 that executes one or more enhancement algorithms. Enhancement techniques are used to artificially increase the quantity and / or type of training data by adding slightly modified synthetic copies of the existing training data or synthetic data newly created from the existing training data. As described herein, inter-scanner and inter-laboratory differences can result in intensity and color variations within digital images. In addition, poor scans can result in gradient variations and blurring effects, assay staining can produce staining artifacts such as background washout, and cell sizes can vary among different tissue / patient samples. These variations and perturbations can have a negative impact on the quality and reliability of deep learning and artificial intelligence networks. The enhancement techniques implemented in the enhancement phase 317 act as regularizers for these variations and perturbations and help reduce overfitting when training machine learning models. It should be understood that the enhancement techniques described herein can be used as regularizers for any number and type of variations and perturbations and are not limited to the various specific examples discussed herein.

[0087] During training period 320, the machine learning algorithm 355 can be trained by the training controller 365 using the labels 345 and the corresponding preprocessed images 340 according to the various workflows described herein. For example, to train the algorithm 355, the preprocessed image 340 can be divided into an image subgroup 340a for training (e.g., 90%) and an image subgroup 340b for validation (e.g., 10%). This division can be done randomly (e.g., 90% / 10% or 70% / 30%), or the division can be done according to more sophisticated validation techniques such as k-fold cross-validation, leave-one-out cross-validation, leave-one-group out cross-validation, nested cross-validation, etc., to minimize sampling bias and overfitting. The division can also be done based on the inclusion of enhanced or synthetic images 352 within the preprocessed image 340. For example, it may be beneficial to limit the number or ratio of synthetic images 352 included within the image subgroup 340a for training. In some cases, the ratio of the original image 335 to the synthetic image 352 is maintained at 1:1, 1:2, 2:1, 1:3, 3:1, 1:4, or 4:1.

[0088] In some cases, the machine learning algorithm 355 includes a CNN, a modified CNN with an encoding layer replaced by a residual neural network (“Resnet”), or a modified CNN with encoding and decoding layers replaced by Resnet. In other cases, the machine learning algorithm 355 can be any suitable machine learning algorithm configured to locate, classify, and / or analyze the preprocessed image 340, such as a two-dimensional CNN (“2DCNN”), Mask R-CNN, U-Net, Feature Pyramid Network (FPN), Dynamic Time Warping (“DTW”) technique, Hidden Markov Model (“HMM”), pure attention-based model, etc., or a combination of one or more of such techniques—for example, Vision Transformer, CNN-HMM, or MCNN (Multi-Scale Convolutional Neural Network). The computing environment 300 can employ the same type of machine learning algorithm or different types of machine learning algorithms trained to detect and classify different cells. For example, the computing environment 300 can include a first machine learning algorithm (e.g., U-Net) for detecting and classifying PD1. The computing environment 500 can also include a second machine learning algorithm (e.g., 2DCNN) for detecting and classifying cluster of differentiation 68 (“CD68”). The computing environment 300 can also include a third machine learning algorithm (e.g., U-Net) for combining the detection and classification of PD1 and CD68. The computing environment 300 can also include a fourth machine learning algorithm (e.g., HMM) for diagnosing a disease for treatment or for prognosticating a subject such as a patient. In other instances in accordance with the present disclosure, other types of machine learning models can also be implemented.

[0089] The training process of the machine learning algorithm 355 includes selecting hyperparameters of the machine learning algorithm 355 from the parameter data store 363, inputting the image subset 340a (e.g., the label 345 and the corresponding preprocessed image 340) into the machine learning algorithm 355, and performing iterative operations to learn a set of parameters (e.g., one or more coefficients and / or weights) of the machine learning algorithm 355. Hyperparameters are settings that can be adjusted or optimized to control the behavior of the machine learning algorithm 355. Most algorithms explicitly define hyperparameters that control different aspects of the algorithm (such as memory or execution cost). However, additional hyperparameters can be defined to adapt the algorithm to a specific scenario. For example, hyperparameters can include the number of hidden units of the algorithm, the learning rate of the algorithm (e.g., 1e-4), the convolutional kernel width, or the number of kernels of the algorithm. In some cases, compared with a typical CNN, the number of model parameters of each convolutional and deconvolutional layer and / or the number of convolutional kernels of each convolutional and deconvolutional layer is reduced by half.

[0090] The image subset 340a can be input into the machine learning algorithm 355 as a batch having a predetermined size. The batch size limits the number of images presented to the machine learning algorithm 355 before parameter updates can be made. Alternatively, the image subgroup 340a can be input into the machine learning algorithm 355 as a time series or sequentially. In either case, in the event that the preprocessed image 340a includes the enhanced or synthetic image 352, the number of original images 335 and the number of synthetic images 352 included in each batch or the manner in which the original images 335 and synthetic images 352 are fed into the algorithm (e.g., every other batch or image is a batch of original images or an original image) can be defined as hyperparameters.

[0091] Each parameter is an adjustable variable such that the value of the parameter is adjusted during training. For example, a cost function or objective function can be configured to optimize the accurate classification of the depicted representation, optimize the characterization of a given type of feature (e.g., characterization of shape, size, uniformity, etc.), optimize the detection of a given type of feature, and / or optimize the accurate localization of a given type of feature. Each iteration can involve learning a set of parameters of the machine learning algorithm 355 that minimizes or maximizes the cost function of the machine learning algorithm 355 such that the value of the cost function using that set of parameters is less than or greater than the value of the cost function using another set of parameters in a previous iteration. The cost function can be configured to measure the difference between the output predicted using the machine learning algorithm 355 and the label 345 included in the training data. For example, for a supervised learning-based model, the goal of training is to learn a function "h()" (sometimes also referred to as a hypothesis function) that maps the training input space X to the target value space Y, h: X → Y, such that h(x) is a good predictor of the corresponding value of y. Various different techniques can be used to learn this hypothesis function. In some techniques, as part of deriving the hypothesis function, a cost or loss function can be defined to measure the difference between the true value of the input and the predicted value of that input. As part of training, techniques such as backpropagation, stochastic feedback, direct feedback alignment (DFA), indirect feedback alignment (IFA), Hebbian learning, etc. are used to minimize this cost or loss function.

[0092] The training iteration continues until a stop condition is met. The training completion condition can be configured to be met when, for example, a predefined number of training iterations have been completed, statistics generated based on testing or validation exceed a predetermined threshold (e.g., a classification accuracy threshold), statistics generated based on a confidence metric (e.g., an average or median confidence metric or a percentage of confidence metrics above a particular value) exceed a predetermined confidence threshold, and / or when a user device participating in the training review closes the training application executed by the training controller 365. Once a set of model parameters has been identified via training, the machine learning algorithm 355 has been trained, and the training controller 365 uses the image subgroup 340b (the test or validation dataset) for an additional process of testing or validation. This validation process can include iterative operations of inputting images from the image subgroup 340b into the machine learning algorithm 355 using validation techniques such as k-fold cross-validation, leave-one-out cross-validation, leave-a-group-out cross-validation, nested cross-validation, etc. to adjust hyperparameters and ultimately find an optimal set of hyperparameters. Once the optimal set of hyperparameters is obtained, a held-out test set of images from the image subgroup 340b is input into the machine learning algorithm 355 to obtain an output, and correlation techniques such as the Bland-Altman method and the Spearman rank correlation coefficient are used to evaluate the relationship between the output and the true values and calculate performance metrics such as error, accuracy, precision, recall, receiver operating characteristic curve (ROC), etc. In some cases, a new training iteration can be initiated in response to receiving a corresponding request or trigger condition from a user device (e.g., initial model development, model update / adaptation, continuous learning, determination of drift within the trained machine learning model 360, etc.).

[0093] It should be understood that other training / validation mechanisms are also contemplated and can be implemented within the computing environment 300. For example, the machine learning algorithm 355 can be trained and hyperparameters can be adjusted on images from the image subgroup 340a, and images from the image subgroup 340b can be used only for testing and evaluating the performance of the machine learning algorithm 355. Additionally, although the training mechanisms described herein focus on training a new machine learning algorithm 355, these training mechanisms can also be used for initial model development, model update / adaptation, and continuous learning of an existing machine learning model 360 trained from other datasets, as described in detail herein. For example, in some cases, the machine learning model 360 may have been pre-disposed using images of other objects or biological structures or images of slices from other subjects or studies (e.g., human trials or murine experiments). In these cases, the machine learning model 360 can be used for initial model development, model update / adaptation, and continuous learning using the preprocessed images 340.

[0094] The trained machine learning model 360 can then be used (during the result generation period 325) to process the new preprocessed image 340 to generate predictions or inferences, such as predicting cell centers and / or location probabilities, classifying cell types, generating cell masks (e.g., per-pixel segmentation masks of the image), predicting the diagnosis of a disease or the prognosis of a subject such as a patient, or combinations thereof. In some cases, the mask identifies the location of the depicted cells associated with one or more biomarkers. For example, given tissue stained for a single biomarker, the trained machine learning model 360 can be configured to: (i) infer the centers and / or locations of the cells, (ii) classify the cells based on characteristics of the staining pattern associated with the biomarker, and (iii) output a cell detection mask for positive cells and a cell detection mask for negative cells. As another example, given tissue stained for two biomarkers, the trained machine learning model 360 can be configured to: (i) infer the centers and / or locations of the cells, (ii) classify the cells based on characteristics of the staining pattern associated with the two biomarkers, and (iii) output a cell detection mask for cells positive for the first biomarker, a cell detection mask for cells negative for the first biomarker, a cell detection mask for cells positive for the second biomarker, and a cell detection mask for cells negative for the second biomarker. As another example, given tissue stained for a single biomarker, the trained machine learning model 360 can be configured to: (i) infer the centers and / or locations of the cells, (ii) classify the cells based on cell characteristics and the staining pattern associated with the biomarker, and (iii) output a cell detection mask for positive cells, a cell detection mask for negative cell codes, and mask cells classified as tissue cells.

[0095] In some cases, the analysis controller 380 generates an analysis result 385 that is used to request an entity to process the underlying image. The analysis result 385 can include a mask superimposed on the new pre-processed image 340 output from the trained machine learning model 360. Additionally or alternatively, the analysis result 385 can include information calculated or determined based on the output of the trained machine learning model, such as a whole-slide tumor score. In an exemplary embodiment, the automated analysis of tissue slides uses an FDA-cleared 510(k)-approved algorithm of the assignee VENTANA. Alternatively or additionally, any other automated algorithm can be used to analyze selected regions of an image (e.g., a masked image) and generate a score. In some implementations, the analysis controller 380 can further respond to instructions from a pathologist, physician, researcher (e.g., related to a clinical trial), subject, healthcare professional, etc. received from a computing device. In some cases, the communication from the computing device includes an identifier for each of a specific group of subjects, which corresponds to a request for an analysis iteration for each subject represented in the group. The computing device can further perform an analysis and / or provide a recommended diagnosis / treatment for a subject based on the output of the machine learning model and / or the analysis controller 380.

[0096] It should be understood that the computing environment 300 is exemplary, and computing environments 300 with different periods and / or using different components are contemplated. For example, in some cases, the network can omit the pre-processing stage 310 such that the images used for training the algorithm and / or the images processed by the model are raw images (e.g., from the image data storage). As another example, it should be understood that each of the pre-processing stage 310 and the training stage 320 can include a controller to perform one or more of the actions described herein. Similarly, while the labeling stage 315 is depicted as being associated with the labeling controller 350 and while the result generation stage 325 is depicted as being associated with the analysis controller 380, the controllers associated with each stage can further or alternatively facilitate other actions described herein in addition to generating labels and / or generating analysis results. As yet another example, for Figure 3The depiction of the computing environment 300 shown lacks the following depicted representations: devices associated with a programmer (e.g., which selects an architecture for a machine learning algorithm 355, defines how various interfaces will operate, etc.); devices associated with a user who provides an initial annotation or annotation review (e.g., during annotation period 315); and devices associated with a user who requests model processing of a given image (which user may be the same or different from the user who provided the initial annotation or annotation review). Although these devices are not depicted, the computing environment 300 may involve the use of one, more than one, or all of the devices, and may in fact involve the use of multiple devices associated with a corresponding plurality of users who provide an initial label or label review and / or multiple devices associated with a corresponding plurality of users who request model processing of various images).

[0097] V. Determine the Consensus Location of Annotations

[0098] When annotations from different annotators are added to an image, various types of annotation conflicts may inevitably occur. For example, one or more annotation-location conflicts between annotations may be found in the image. An annotation-location conflict can be identified when: (i) one annotation identifies a cell as being located at a particular location in the image, but one or more other annotations identify the same cell as being located at a different location in the image; (ii) one annotation identifies that there are multiple cells located at a given location, but one or more other annotations identify that there is a single cell located at the same given location; or (iii) one annotation identifies that there is a cell located at a certain location, but one or more other annotations identify that there is no cell located at that location. Thus, determining the consensus location between annotations can ensure that the corresponding labels point to the same location and remove annotations that may not accurately identify cells in the image (e.g., duplicate annotations).

[0099] A. Input Data

[0100] To receive and process annotations to be associated with a digital pathology image (or FOV of an image), an annotation processing system (e.g., Figure 1 annotation processing system 102 of Figure 3The annotation processing system 351) parses one or more schema files (or other types of data structures) provided by a set of annotators. The schema file may include a set of annotations, where each annotation identifies an initial position of a specific image region of the image and an initial label for classifying a biological structure depicted at the initial position of the image (e.g., a stained tumor cell). Annotations can be added by the annotator using an annotation tool (e.g., dpath). A set of protocols can be defined for the annotator such that annotations can be generated based on the same configuration of the image. The instructions can include preconfigured values for the following items: (i) magnification level; (ii) type of cells to be annotated; (iii) definition of each label; (iv) additional metrics for ambiguous cases; (v) use of a specific color profile; (vi) seed placement of annotations; and (iv) indication of cells at the boundaries of the image.

[0101] In addition to the above information, the annotation can also be associated with a position within the image. In some cases, the annotation includes x and y coordinates identifying its position relative to the image. Additional data can be added to the annotation based on the input provided by the annotator. Annotations from the annotator can be exported as a schema file, such as an XML file. In some cases, a single schema file is generated for the entire image. Additionally or alternatively, a schema file can be generated for each FOV of the image.

[0102] The annotation processing system can receive and parse the schema file to determine the corresponding image (or FOV of the image) for which annotations should be added and associate each annotation with the corresponding position in the image. The image with annotations can be processed as part of a training dataset for training a machine learning model. In some cases, the image is processed as part of a validation dataset for validating the machine learning model. Additionally or alternatively, the image can be processed as part of a test dataset for testing the machine learning model.

[0103] B. Identifying Annotation Clusters

[0104] Once the annotations are processed, the annotation processing system determines which annotations submitted by the individual annotators relate to the same position. Figure 4 A schematic diagram showing a process 400 for identifying clusters of an annotation set for a slide image according to some embodiments is shown. Process 400 can traverse the annotations submitted by each annotator to identify clusters, thereby associating the annotations as corresponding to the same region.

[0105] As Figure 4As shown, multiple annotations associated with an image can include a first set of annotations 402 associated with a first annotator, "Pathologist 1". Each annotation in the first set 402 includes an identifier (e.g., "p11") and a set of x and y coordinates that identify the location of the annotation within the image. The annotations of the first set of annotations 402 can be compared with the annotations from the remaining set of annotations 404. The remaining set of annotations 404 can include a second set of annotations associated with a second annotator, "Pathologist 2", and a third set of annotations associated with a third annotator, "Pathologist 3", where the annotations of the second and third sets include respective identifiers (e.g., "p21") and respective sets of x and y coordinates.

[0106] Based on the above information, the annotation processing system can perform clustering of the multiple annotations based on the respective x and y coordinates of the multiple annotations to form one or more annotation clusters. In some cases, one or more clusters are formed by performing a distance-based search among the multiple annotations. The cluster can be defined as a set of annotations within a predetermined distance. In some cases, a threshold is used to determine the set of annotations, where the threshold identifies the number of pixels between the annotations (e.g., 10 pixels). The threshold can be adjusted by the annotator, which may increase or decrease the number of annotations to be included in the cluster. By adjusting the threshold to increase or decrease the number of annotations, the size of the cluster in the image can be defined.

[0107] As Figure 4 shown, the first cluster set 406 includes a first cluster (shown in pink) that includes the "p11" annotation from the first set of annotations 404 as well as the "p22" and "p31" annotations. The "p11" annotation can include the "p22" and "p31" identifiers ("Cp-id") as the nearest points. The first cluster can then be identified by the associate key of the "p11" of the first cluster for the "p11", "p22", and "p31" annotations. Similarly, the first cluster set 406 can include a second cluster (shown in teal) that includes the "p12", "p21", and "p32" annotations, where the second cluster can be identified by its associate key, "p12".

[0108] Additional cluster combinations can be formed by repeating the above process based on other sets of annotations (e.g., a second set of annotations and a third set of annotations). For example, a second set of annotations 408 associated with a second annotator "Pathologist 2" can be compared with another remaining set of annotations 410. The other remaining set of annotations includes a first set of annotations corresponding to "Pathologist 1" and a third set of annotations corresponding to "Pathologist 3". Clustering can be performed to generate a second annotation cluster 412, which can include an annotation "p23" of the second set of annotations 408 associated with "p32" of the third set of annotations. In addition, "p22" and "p31" are associated with the same cluster based on an association key "p11". In some cases, clustering is iterated through the respective sets of annotations until all possible combinations of annotation clusters are formed.

[0109] Then, the first cluster 406 and the second cluster 412 can be merged to generate a merged annotation cluster 414. The merged cluster can be generated by merging the clusters of the first cluster set 406 with the corresponding clusters of the second cluster set 412. For example, the "p32" annotation can be identified as a closet annotation of the "p12" annotation of the first set of annotations 404 and the "p23" annotation of the second set of annotations 408. Based on this association, the "p23" annotation can be merged into the cluster with the association key "p12". In another example, the "p33" annotation can be identified as a closet annotation of the "p21" annotation, which can also be identified as the closest annotation to the "p12" annotation. Thus, the "p33" annotation can also be merged into the cluster with the association key "p12".

[0110] Figure 5 An exemplary image 500 including the locations of annotations is shown according to some embodiments. Each depicted cell in the image 500 is associated with one or more colored circles (e.g., pink, orange, blue, green, red), where each circle represents an annotation 502. The color of the circle identifies the annotator who generated the annotation. For example, the red circle identifies the first pathologist who submitted the annotation for the image 500, and the blue circle identifies the second pathologist who submitted the annotation for the image 500. Annotations within a predetermined distance threshold (e.g., 10 pixels) can be associated as part of a cluster 504. The cluster 504 can represent a single cell and can be identified by a yellow line or a black line connecting the circles. In some cases, two or more clusters are merged into a single merged cluster 506.

[0111] C. Removing Annotation Conflicts

[0112] Annotation - location conflicts can occur when generating clusters for annotations. When two or more annotations (representing multiple cells) from a particular annotator are identified as involving the same cluster (representing a single cell) during a distance - based search, an annotation - location conflict may occur. An annotation - location conflict can occur when two cells are very close to each other in a given image. In such cases, it is difficult to determine which cell a particular annotation refers to. In another instance, a pathologist may incorrectly indicate that there are two tumor cells when there is actually a single tumor cell, resulting in an annotation - location conflict.

[0113] Figure 6 An exemplary set of annotations identified in a training slide image 600 according to some embodiments is shown. Each depicted cell in image 600 is associated with one or more colored circles (e.g., pink, orange, blue, green, red), where each circle represents an annotation. The color of the circle identifies the annotator who submitted the annotation. The circles in image 600 are connected by black or yellow lines that represent a certain cluster. A cluster may represent or identify a single cell.

[0114] Within a single cluster, two or more annotations (e.g., multiple cells) can be identified from the same annotator. For example, cluster 602 can be identified by at least three orange circles, where the corresponding annotator views the image region as depicting three cells. Based on the three annotations from the same annotator, the annotation processing system can determine that cluster 602 includes at least one annotation - location conflict because it is assumed that cluster 602 represents only a single cell but currently includes annotations indicating multiple cells. The annotation - location conflict can be indicated by the annotation processing system by marking cluster 602 (e.g., the association key of the cluster) as having "conflict" and displaying such a cluster on a graphical user interface.

[0115] In some cases, the annotation processing system indicates an annotation - location conflict by providing a visual marker on image 600. For example, the first cluster set 604 can be indicated as not having an annotation - location conflict based on one or more "black" lines connecting the annotations. In contrast, the second cluster set 606 can be indicated as having one or more annotation - location conflicts based on one or more "yellow" lines connecting the annotations.

[0116] Two types of annotation - location conflicts may be encountered during this process. The two types of annotation - location conflicts may include: (i) a first conflict type that involves the location of two or more annotations from a single annotator; and (ii) a second conflict type that involves the location of two or more annotations from each of two or more annotators.

[0117] Figure 7Displays an exemplary image showing a process 700 for resolving annotation conflicts according to some embodiments. Process 700 may include removing annotation conflicts for a first type of conflict. The preprocessed image 702 shows an exemplary image including one or more annotation conflicts, which includes a cluster 704. For each of the two or more annotations from the same annotator, the annotation processing system may determine a statistical value (e.g., mean, median) representing the distance from the annotation to other annotations within the corresponding cluster. The statistical values between the two or more annotations from the same annotator may be compared. The annotation with the lowest statistical value may be determined to be part of the cluster, while the remaining annotations of the conflict may be removed from being part of the cluster. The postprocessed image 706 shows another exemplary image in which the one or more annotation conflicts have been removed. As shown in the postprocessed image 706, the yellow line connecting the orange circle 708 to the cluster 704 has been removed. In some cases, the orange circle 708 may form its own cluster and identify different objects or regions within the image.

[0118] Figure 8 Displays another set of exemplary images showing a process 800 for removing annotation conflicts according to some embodiments. The preprocessed image 802 shows an exemplary image including one or more annotation conflicts, which includes a merged cluster 804. Process 800 may include removing annotation conflicts for a second type of conflict. The annotation processing system applies a clustering algorithm to the cluster that includes two or more annotations from each of two or more annotators, such that the cluster can be divided into sub-clusters. In some cases, the clustering algorithm includes the k-means algorithm, which divides the cluster into a "k" number of sub-clusters. The "k" number of sub-clusters may correspond to the maximum number of annotations input by a particular annotator for a given cluster.

[0119] For each sub-cluster that includes the two or more annotations from the same annotator, the annotation processing system may determine, for each of the two or more annotations, a statistical value (e.g., mean, median) representing the distance from the annotation to other annotations within the corresponding cluster. The statistical values between the two or more annotations from the same annotator may be compared. The annotation with the lowest statistical value may be determined to be part of the sub-cluster, while the remaining annotations may be removed from being part of the sub-cluster. The postprocessed image 806 shows another exemplary image in which the one or more annotation conflicts have been removed. Sub-clusters 808 and 810 show two clusters divided from the merged cluster 804, where the yellow line connecting the sub-clusters 808 and 810 is removed.

[0120] D. Determining a consensus position

[0121] After resolving annotation conflicts, the annotation processing system can determine clusters that can be retained in the image. The clusters in the cluster set can include a certain number of annotations that exceed a confidence threshold. The confidence threshold can correspond to the minimum count of annotations that need to be present within a given cluster. For example, the confidence threshold can specify that at least three annotations need to be present for a cluster to be retained in the image. A cluster that includes four annotations can be determined to exceed the confidence threshold and can be added to the cluster set that can be retained in the image. Conversely, a cluster that includes two annotations can be determined to not exceed the confidence threshold. Such clusters are discarded from the image, and the annotations of the clusters are also discarded from the image. Thus, the cluster set with a sufficient number of annotations can then be further processed. In some cases, the confidence threshold is a percentage value (e.g., 0.5, 40%) compared to a statistical value that represents the portion of annotations in the cluster that exceed the total number of annotators who generated annotations for the image. The confidence threshold can be adjusted based on user input, which may result in an increase or decrease in the true labels present in the image.

[0122] After determining clusters that exceed the confidence threshold, a consensus position can be determined for each cluster identified in the image. Each cluster can include a set of annotations submitted by multiple annotators, and the consensus position of the cluster can be determined based on the median coordinates among the positions of the set of annotations associated with the cluster. For example, the x and y coordinates of each annotation of a given cluster can be identified. The median x coordinate can be determined from the x coordinates of the set of annotations of the cluster. Similarly, the median y coordinate can be determined from the y coordinates of the set of annotations of the cluster. The median x coordinate and the median y coordinate can be used to determine the consensus position of the cluster such that the true label can be located at the consensus position of the image. Additionally or alternatively, other statistical values such as the average can be used instead of the median x - y coordinates.

[0123] In some cases, the x and y coordinates of each annotation are assigned corresponding weight values based on the annotator who submitted the annotation. For example, the x and y coordinates of an annotation submitted by a senior pathologist can be assigned a weight value that is greater than another weight value assigned to the x and y coordinates of another annotation submitted by a less - senior pathologist. Thus, the consensus position can be determined based on the weighted x and y coordinates of the set of annotations within a given cluster. The consensus position of the cluster can be assigned a true label that identifies the target type (e.g., tumor cells) of the biological structure corresponding to the object or region within the image. Moreover, the true label at the consensus position can be used to train a machine - learning model that predicts the type of biological structure in the corresponding region of other slide images.

[0124] E. Method for Identifying Annotation - Location Conflicts in Images

[0125] Figure 9Illustrates process 900 for identifying annotation-location conflicts in an image. For illustrative purposes, reference is made to the components shown in Figure 1 and / or Figure 3 in describing process 900, although other implementations are possible. For example, Figure 1 the program code of annotation processing system 102 and / or Figure 3 the program code of annotation processing system 351 (which is stored in a non-transitory computer-readable medium) is executed by one or more processing devices to cause the server system to perform one or more operations described herein.

[0126] At step 902, the annotation processing system accesses a plurality of annotations associated with an image depicting at least a portion of a biological sample. The image can be a digital pathology image depicting at least a portion of a biological sample (e.g., tissue) obtained from a subject. The plurality of annotations can be used to generate ground truth labels for supervised training, validation, and / or testing of a machine learning model. Each annotation in the plurality of annotations can be located at a position within the image that identifies an object or region within the image. In some cases, each annotation includes an identifier associated with the annotator (e.g., pathologist) who generated the annotation.

[0127] At step 904, the annotation processing system identifies a first annotation generated by a first annotator and a second annotation generated by a second annotator from the plurality of annotations. In some cases, the first annotation further includes a label (e.g., stained tumor cells, unstained tumor cells, normal cells) that classifies the target type of the biological structure depicted at a first position in the image. Similarly, the second annotation includes its corresponding label that classifies the same or different target type of the biological structure depicted at a second position in the image.

[0128] At step 906, the annotation processing system determines a first distance value between a first position of the first annotation and a second position of the second annotation.

[0129] At step 908, the annotation processing system determines that the first distance value is below a predetermined threshold. The predetermined threshold can be used to determine whether the first annotation and the second annotation correspond to the same object or region within the image. In some cases, the predetermined threshold is adjusted to increase or decrease the number of annotations that commonly identify the same object or region of the image.

[0130] At step 910, in response to determining that the first distance value is below the predetermined threshold, the annotation processing system determines that both the first annotation and the second annotation identify a first object or region within the image.

[0131] At step 912, the annotation processing system identifies a third annotation generated by a first annotator from the plurality of annotations, wherein the third annotation is located at a third location different from the first location.

[0132] At step 914, the annotation processing system determines that a second distance value between the second location of the second annotation and the third location of the third annotation is below a predetermined threshold.

[0133] At step 916, in response to determining that the second distance value is below the predetermined threshold, the annotation processing system determines that the third annotation also identifies a first object or region within the image.

[0134] At step 918, in response to determining that the third annotation identifies a first object or region within the image, the annotation processing system generates an output indicating that there is an annotation conflict for the object or region within the image. Thereafter, process 900 terminates.

[0135] F. Method for Removing Annotation - Location Conflicts in an Image

[0136] Figure 10 Process 1000 for removing annotation - location conflicts and determining a consensus label in an image is shown in accordance with some embodiments. For illustrative purposes, process 1000 is described with reference to the components shown in Figure 1 and / or Figure 3 although other implementations are possible. For example, Figure 1 the program code of the annotation processing system 102 and / or Figure 3 the annotation processing system 351 (which is stored in a non - transitory computer - readable medium) is executed by one or more processing devices to cause the server system to perform one or more of the operations described herein.

[0137] At step 1002, the annotation processing system identifies a location conflict among a first annotation, a second annotation, and a third annotation of the image. The training image can be a digital pathology image depicting at least a portion of a biological sample (e.g., tissue) obtained from a subject. The first annotation, the second annotation, and the third annotation can be used to generate ground truth labels for supervised training, validation, and / or testing of a machine - learning model. The first annotation and the third annotation are generated by a first annotator, and the second annotation is generated by a second annotator. The first annotation, the second annotation, and the third annotation can be identified as identifying a first object or region within the image. The annotation conflict can be determined using the steps of Figure 9 process 900.

[0138] At step 1004, the annotation processing system compares a first distance value between the first annotation and the second annotation with a second distance value between the second annotation and the third annotation. The first distance value may correspond to the distance between a first position of the first annotation and a second position of the second annotation. The second distance value may correspond to the distance between the second position and a third position of the third annotation.

[0139] At step 1008, if the first distance value is greater than the second distance value (the "yes" path from step 1006), the annotation processing system may resolve the annotation-position conflict by determining the third annotation as corresponding to a second object or region (e.g., a different cluster) of the image. At step 1010, if the second distance value is greater than the first distance value (the "no" path from step 1006), the annotation processing system may resolve the annotation-position conflict by determining the first annotation as corresponding to a second object or region (e.g., a different cluster) of the image.

[0140] At step 1012, the annotation processing system determines a consensus position of the first object or region of the image based on the remaining annotations associated with the first object or region. The consensus position may be determined based on the median coordinates among the positions of the remaining annotations associated with the first object or region of the image. Additionally or alternatively, other statistical values such as an average value may be used instead of the median x-y coordinates.

[0141] In some cases, the annotation processing system resolves another type of annotation-position conflict involving two or more annotations from each of two or more annotators. For example, the annotation processing system identifies a fourth annotation generated by a second annotator. The fourth annotation is located at a fourth position of the image that is different from the second position of the second annotation. Then, the annotation processing system determines a third distance value between the third position and the fourth position, and further determines that the third distance value is below a predetermined threshold. In response to determining that the third distance value is below the predetermined threshold, the annotation processing system determines that the first annotation, the second annotation, the third annotation, and the fourth annotation identify the same object or region within the image. Since two or more annotations from each of the first annotator and the second annotator are identified for the same region, the annotation processing system may determine that a different type of annotation-position conflict has occurred.

[0142] To resolve the other annotation-position conflict, the annotation processing system may apply a clustering algorithm to the first position, the second position, the third position, and the fourth position. The clustering algorithm may be applied to determine: (i) the first annotation and the second annotation correspond to a first sub-region of the first object or region; and (ii) the third annotation and the fourth annotation correspond to a second sub-region of the first object or region. In some cases, the clustering algorithm is a k-means clustering algorithm.

[0143] At step 1014, the annotation processing system assigns a ground truth label for the consensus location for the image. The ground truth label identifies the target type (e.g., tumor cell) of the biological structure of the first object or region within the image. The ground truth label at the consensus location can be used to train a machine learning model that predicts the type of biological structure in the corresponding region of other slide images. Using the ground truth label with the consensus location to train the machine learning model can correspond to Figure 3 the training process described in. For different types of annotation-location conflicts that result in a first sub-region and a second sub-region, the annotation processing system can assign corresponding ground truth labels for each of the first sub-region and the second sub-region. Thereafter, process 1000 terminates.

[0144] VI. Determine the consensus label for the annotation

[0145] After determining the consensus location for the annotation set, the annotation processing system can determine the ground truth label for the consensus location for the image, where the ground truth label can be used for training, validation, and / or testing of the machine learning model. In some cases, if all the annotations in the annotation set indicate the same target type (e.g., tumor cell) of the biological structure, the ground truth label is determined. Similarly, the ground truth label can be determined based on the vast majority of the annotation set that identifies the specific target type of the biological structure. However, in some cases, there is no clear consensus among the annotation sets regarding the target type of the biological structure. For example, the consensus location can include four annotations, where two annotators mark the location as having stained tumor cells, while the other two annotators mark the location as having unstained tumor cells.

[0146] In such cases of annotation-label conflicts, the annotation-processing system can determine the consensus label for the object or region of the image. In particular, each annotation in the annotation set can be assigned a corresponding weight value. The weight value is determined based on the performance and historical data associated with the annotator who generated the annotation. For example, the annotator of the first annotation can be a senior pathologist with several years of experience. The weight of the first annotation can be assigned a value that is relatively greater than another weight of the second annotation associated with a less experienced pathologist.

[0147] Then, each target type of the biological structure can be represented by an aggregated weight that is derived from the weights assigned to the corresponding annotations. In response to determining that a specific target type of the biological structure has the highest aggregated weight, a ground truth label can be generated to include the specific target type of the biological structure. Thus, a consensus label can be determined for the region or object of the image, where the ground truth label can include information from the consensus label. Even when some annotations at the same location provide conflicting information, the determination of the consensus label can facilitate the accurate training, validation, and testing of the machine learning model.

[0148] A. Determine whether the set of annotations can be used to determine a consensus label for a region of an image

[0149] An annotation processing system can determine whether a set of annotations corresponding to a particular region or object of an image (e.g., a consensus location) can be used to generate a consensus label.

[0150] Figure 11 A schematic diagram illustrating a process 1100 for removing annotation-label conflicts and determining a consensus label for an image according to some embodiments is shown. At block 1102, an annotation processing system receives a set of annotations corresponding to a region or object of an image. The image can be part of a training data set for training a machine learning model. In some cases, the image is part of a validation data set for validating a machine learning model. Additionally or alternatively, the image can be part of a test data set for testing a machine learning model. The region or object can be identified based on a consensus location determined using the processes described in Figure 9 and Figure 10 The weight associated with the set of annotations can be retrieved by the annotation processing system at block 1104. For each annotation in the set of annotations, a weight can be assigned to the annotation, where the weight can be determined based on the performance and historical data associated with the annotator who generated the annotation. For example, the weight can be determined based on the level of expertise of the annotator who generated the annotation. In some cases, each weight in the weights corresponds to a value between 0 and 1, and the sum of the weights of all annotators can add up to 1.

[0151] In some cases, the annotation processing system calculates the sum of the weights and compares the sum to a predetermined confidence threshold. The confidence threshold is a user-adjustable value that can specify the number of annotations required to generate a consensus label. Comparing the weights to the confidence threshold can help determine whether the set of annotations contains sufficient information to generate a consensus label.

[0152] As an illustrative example, two annotations corresponding to a region of an image can be determined to be sufficient to generate a consensus label based on the sum of the corresponding weights of two annotations that exceed the confidence threshold. In another example, four annotations corresponding to another region of an image can be determined to be insufficient to generate a consensus label based on the sum of their corresponding weights that do not exceed the confidence threshold. In some cases, the annotation processing system includes one or more other annotations corresponding to locations near a particular region or object, such that some annotations with higher weights can also be considered.

[0153]

[0154] ​At block 1106, the annotation processing system can determine whether the set of annotations for a particular region or object is sufficient to generate a consensus label by comparing the count of the set of annotations for that region or object with a second confidence threshold. For example, the second confidence threshold can be three, where if the count of the set of annotations is equal to or greater than three, the set of annotations can be considered sufficient to generate a consensus label. If the count of the set of annotations does not exceed the second confidence threshold, the annotation processing system can discard the set of annotations from the particular region or object such that the set of annotations is no longer processed for generating a ground truth label. Conversely, if the count of the set of annotations exceeds the second confidence threshold, the annotation processing system can use the set of annotations for generating a consensus label. This process can be repeated for each region or object in the image. Thus, the set of annotations across different regions or objects can then be used to generate a consensus label for the image.

[0155] B. Resolving label conflicts in the set of annotations

[0156] For the set of annotations corresponding to an object or region of an image, the annotation processing system can remove any label conflicts and generate a consensus label for the object or region. Recall Figure 11 , the annotation processing system can identify the target type of the biological structure (block 1108) based on the weights assigned to the set of annotations. For each target type of the biological structure identified in the set of annotations, an aggregated weight can be determined based on a subset of the annotations that includes the target type of the biological structure. For example, a first aggregated weight for a target stained tumor cell can be based on the weights of three annotations in the set of annotations that identify the corresponding region of the image as depicting a tumor cell. In another example, a second aggregated weight for a target immune cell (e.g., lymphocyte) can be based on the weights of two remaining annotations in the set of annotations that identify the corresponding region of the image as depicting an immune cell. In some cases, the aggregated weight corresponds to the sum of all the weights associated with the corresponding target type of the biological structure. Additionally or alternatively, the aggregated weight can correspond to the average or median of all the weights associated with the corresponding target type of the biological structure.

[0157] The annotation processing system can compare the aggregated weights to determine the target type of the biological structure with the highest weight. The target type of the biological structure with the highest aggregated weight can be identified as the consensus label for the region or object of the image. Thus, a ground truth label including the target type can then be generated for the image. Continuing from the above example, even when the target stained tumor cell is associated with a larger number of annotations, the second aggregated weight for the target immune cell can be determined to be greater than the first aggregated weight. In response to this determination, the annotation processing system generates a ground truth label for the corresponding region of the image, where the ground truth label includes immune cell.

[0158] Figure 12 An exemplary image 1200 including a consensus label and a location according to some embodiments is shown. Each point represents a consensus location depicting a biological structure (e.g., a tumor cell). The color of the point represents the ground truth label for the location. For example, a "blue" point represents other types of cells (e.g., lymphocytes, fibroblasts), a "red" point represents a stained tumor cell, and a "green" point represents an unstained tumor cell.

[0159] In some cases, for a corresponding region of the image, there are cases where at least two aggregation weights are equal to each other. Recall Figure 11 , the annotation processing system can determine a secondary weight for each annotation in the annotation set (block 1110). The secondary weight can correspond to a value assigned based on an accuracy metric assigned to the annotator who generated the annotation. The accuracy metric can be determined based on the number of other annotations of other regions of the image that match the corresponding consensus label provided by the annotator. The secondary weight can be used to normalize the aggregation weight, where the target type of the biological structure associated with the larger normalized weight can be identified as the consensus label for the corresponding region.

[0160] As an illustrative example, an object or region of an image includes two annotations, where the first annotation indicates a stained tumor cell and the second annotation indicates an unstained tumor cell. The annotation processing system determines that the two annotations have equal weights assigned to them, thus identifying a label conflict between the annotations. In response to such a determination, the secondary weight of the first annotation can be calculated based on determining that 90% of the annotations generated by the annotator for other regions match the corresponding consensus label. Then, the secondary weight of the second annotation can be calculated based on determining that 80% of the annotations generated by another annotator for other regions match the corresponding consensus label, where the secondary weight of the second annotation is less than the secondary weight of the first annotation based on their respective accuracy metrics. Based on the comparison, the stained tumor cell indicated by the first annotation can be selected to generate the ground truth label for the object or region of the image. Thus, by using weights and secondary weights, the annotation processing system can resolve annotation-label conflicts in various regions of the image.

[0161] C. Adjusting Weights Based on Annotator Performance

[0162] The consensus label can be used to train a machine learning model, but can also be used to re-evaluate an annotator and adjust the corresponding weight of the annotator. Refer to Figure 11, the annotation processing system can adjust the weight associated with the annotator by tracking the performance of the annotator in different images (e.g., training images, validation images) or different FOVs within the same image (block 1112). For an annotator who generates one or more annotations in an image (or FOV in the image), the annotation processing system can determine the number of annotations that match the corresponding consensus label in the given image. The number of annotations can be compared with the total number of annotations generated by the annotator for the image. Based on the comparison, the annotation processing system can adjust the weight associated with the annotator. The weight adjustment process can continue across other images to fine-tune the weight of the annotator. The above process can be repeated across other annotators who generate annotations for the image. Additionally or alternatively, the annotation processing system can adjust the weight of the annotator based on the number of annotations that match the corresponding consensus label for each target type of the biological structure (the target type is defined for the image).

[0163] In some cases, the weight of the annotator is adjusted based on the performance of the annotator relative to the performance of other annotators for the image. For example, it can be determined that the first annotator has 85% of the annotations that match the corresponding consensus label of the image. It can be determined that the second annotator has 89% of the annotations that match the corresponding consensus label of the image. The performance of the first annotator and the second annotator can be compared, and based on the comparison, a higher weight can be assigned to the second annotator than the first annotator (e.g., the assigned weight is proportional to the percentage of annotations that match the corresponding consensus label). If there are three or more annotators who generate annotations for the image, the annotators can be ranked based on their performance for the image, and the corresponding weight of the annotator can be adjusted based on the ranking.

[0164] In fact, when performing the fine-tuning step across multiple FOVs (e.g., 100 FOVs) or multiple images, the weight can accurately reflect the expertise and performance of the annotator. In some cases, a particular annotator assigned a low weight is retrained and then allowed to provide annotations for other images. Additionally or alternatively, the annotations of a particular annotator assigned a low weight can be removed from use during the process to generate the ground truth labels for the image.

[0165] D. Method for Removing Annotation-Label Conflicts in Images

[0166] Figure 13 Process 1300 for removing annotation-label conflicts and determining the consensus label in an image according to some embodiments is shown. For illustrative purposes, process 1300 is described with reference to the components shown in Figure 1 and / or Figure 3 although other implementations are also possible. For example, Figure 1The annotation processing system 102 and / or Figure 3 The program code of the annotation processing system 351 (which is stored in a non-transitory computer-readable medium) is executed by one or more processing devices to cause the server system to perform one or more operations described herein.

[0167] At step 1302, the annotation processing system receives a plurality of annotations corresponding to a first object or region of an image. The training image can be a digital pathology image depicting at least a portion of a biological sample (e.g., tissue) obtained from a subject. The plurality of annotations can be used to generate ground truth labels for supervised training, validation, and / or testing of a machine learning model. Each annotation in the annotation set identifies a target type of a biological structure depicted in the first object or region. For example, target types of biological structures include tumor cells, immune cells, etc. In some cases, each annotation includes an identifier associated with the annotator who generated the annotation.

[0168] At step 1304, the annotation processing system identifies a first annotation generated by a first annotator from the plurality of annotations. The first annotation can identify a first target type of a biological structure (e.g., tumor cells). In some cases, the first annotator is associated with a first weight. For example, a senior pathologist with many years of experience can be assigned a weight value that is greater than the weight value assigned to a less experienced junior pathologist.

[0169] At step 1306, the annotation processing system identifies a second annotation generated by a second annotator from the plurality of annotations. The second annotation identifies a second target type of a biological structure (e.g., immune cells), where the second target type is different from the first target type. Thus, the first annotation and the second annotation correspond to the same object or region, but indicate different target types of biological structures. In some cases, the second annotator is associated with a second weight, which can be the same as or different from the first weight assigned to the first annotator.

[0170] At step 1308, the annotation processing system determines that the first weight is greater than the second weight. Referring back to the above example, the first annotator can be a senior pathologist who has been assigned a first weight that is greater than the second weight assigned to the second annotator, who can be a junior pathologist.

[0171] At step 1310, in response to determining that the first weight is greater than the second weight, the annotation processing system generates a ground truth label for the first object or region of the image. The ground truth label can identify the first target type of the biological structure. Continuing from the above example, the ground truth label can include the "tumor cell" label. The ground truth label at the consensus location can be used to train a machine learning model that predicts the type of biological structure in the corresponding region of other slide images. Using the ground truth label with the consensus label and location to train the machine learning model can correspond to Figure 3 the training process described in

[0172] In some cases, the annotation processing system evaluates whether a ground truth label can be generated for the first object or region based on a confidence threshold. For example, the annotation processing system can determine a statistical value (e.g., average, sum) based on the first weight and the second weight, and determine whether the statistical value exceeds the confidence threshold. In some cases, the confidence threshold includes a value that can be adjusted by one or more users. In response to determining that the statistical value does not exceed the confidence threshold, the annotation processing system can generate an output indicating that no consensus has been reached between the first annotator and the second annotator for the ground truth label. Thus, the comparison of the weights with the confidence threshold can filter out annotations with low weighted values, which may indicate that the consensus label is not accurate enough.

[0173] In some cases, the annotation processing system utilizes the assigned weights of the annotators to remove annotation-label conflicts that occur in three or more annotations. The annotation processing system identifies a third annotation generated by a third annotator from the plurality of annotations. The third annotation identifies a second target type of the biological structure (e.g., immune cell), and the third annotator is associated with a third weight. The annotation processing system generates another statistical value based on the second weight and the third weight. Another statistical value is generated based on the second weight and the third weight because both the second annotation and the third annotation identify the same target type of the biological structure. In some cases, another statistical value corresponds to the sum of the second weight and the third weight. The annotation processing system determines that another statistical value is greater than the first weight, where the annotation processing system can redefine the ground truth label for the first object or region of the image to identify the second target type of the biological structure. Thereafter, process 1300 terminates.

[0174] VII. Other Considerations

[0175] Some embodiments of the present disclosure include a system that includes one or more data processors. In some embodiments, the system includes a non-transitory computer-readable storage medium containing instructions that, when executed on the one or more data processors, cause the one or more data processors to perform some or all of one or more methods disclosed herein and / or some or all of one or more processes disclosed herein. Some embodiments of the present disclosure include a computer program product tangibly embodied in a non-transitory machine-readable storage medium, which includes instructions configured to cause one or more data processors to perform some or all of one or more methods disclosed herein and / or some or all of one or more processes disclosed herein.

[0176] The terms and expressions that have been employed are used in a descriptive rather than a restrictive sense, and in using such terms and expressions, there is no intention to exclude any equivalents of the features shown and described or portions thereof, but it should be recognized that various modifications are possible within the scope of the invention as claimed. Accordingly, it should be understood that although the invention as claimed has been specifically disclosed by way of embodiments and optional features, those skilled in the art may adopt modifications and variations of the concepts disclosed herein, and such modifications and variations are considered to be within the scope of the invention as defined by the appended claims.

[0177] The following description provides only preferred exemplary embodiments and is not intended to limit the scope, applicability, or configuration of the present disclosure. On the contrary, the following description of the preferred exemplary embodiments will provide those skilled in the art with a viable description for implementing various embodiments. It should be understood that various changes may be made to the functions and arrangements of the elements without departing from the spirit and scope set forth in the appended claims.

[0178] Specific details are given in the following description to provide a thorough understanding of the embodiments. However, it should be understood that the embodiments may be practiced without these specific details. For example, circuits, systems, networks, processes, and other components may be shown in block diagram form as components to avoid obscuring the embodiments in unnecessary detail. In other instances, well-known circuits, processes, algorithms, structures, and techniques may be shown without unnecessary detail to avoid obscuring the embodiments.

Claims

1. A method, comprising: Accessing a plurality of annotations associated with the image at positions identifying objects or regions within the identification image, wherein the plurality of annotations are used to generate ground truth labels for supervised training, validation, and / or testing of a machine learning model, wherein each of the plurality of annotations is located at a position within a training image identifying a position of an object or region within the training image, and wherein each annotation includes an identifier associated with the annotator who generated the annotation; Identifying, from the plurality of annotations, a first annotation generated by a first annotator and a second annotation generated by a second annotator; Determining a first distance value between a first position of the first annotation and a second position of the second annotation; Determining that the first distance value is below a predetermined threshold; In response to determining that the first distance value is below the predetermined threshold, determining that both the first annotation and the second annotation identify a first object or region within the image; Identifying, from the plurality of annotations, a third annotation generated by the first annotator, wherein the third annotation is located at a third position different from the first position; Determining that a second distance value between the second position of the second annotation and the third position of the third annotation is below the predetermined threshold; In response to determining that the second distance value is below the predetermined threshold, determining that the third annotation also identifies the first object or region within the image; And In response to determining that the third annotation identifies the first object or region within the image, generating an output indicating that there is an annotation conflict for the object or region within the image.

2. The method according to claim 1, further comprising: Determining that the first distance value is less than the second distance value; Determining that the third annotation identifies a second object or region within the image; And Generating another output indicating that the annotation conflict has been resolved.

3. The method according to claim 1 or claim 2, further comprising: Identifying, from the plurality of annotations, a fourth annotation generated by the second annotator, wherein the fourth annotation is located at a fourth position within the image different from the second position of the second annotation; Determining a third distance value between the third position and the fourth position; Determining that the third distance value is below the predetermined threshold; In response to determining that the third distance value is below the predetermined threshold, determining that both the third annotation and the fourth annotation identify the first object or region within the image; Applying a clustering algorithm to the first position, the second position, the third position, and the fourth position to determine: (i) the first annotation and the second annotation correspond to a first sub-region of the first object or region; and (ii) the third annotation and the fourth annotation correspond to a second sub-region of the first object or region.

4. The method according to claim 3, wherein the clustering algorithm comprises a k-means clustering algorithm.

5. The method according to any one of claims 1 to 4, wherein the first object or region depicts a biological structure.

6. The method according to claim 5, wherein: The first annotation further identifies the target type of the biological structure depicted in the first object or region; and The second annotation further identifies the target type of the biological structure depicted in the first object or region.

7. The method according to claim 6, further comprising: Generating a ground truth label for the first object or region within the image, wherein the ground truth label includes the target type of the biological structure; And Providing the image with the ground truth label for supervised training of the machine learning model to predict whether another biological structure depicted in another image corresponds to the target type of the biological structure.

8. The method according to claim 6, wherein the target type of the biological structure corresponds to stained tumor cells, unstained tumor cells, or normal cells.

9. A method, comprising: Receiving a plurality of annotations corresponding to a first object or region of an image, wherein the image depicts at least a portion of a tissue sample, wherein the plurality of annotations are used to generate ground truth labels for supervised training, validation, and / or testing of a machine learning model, wherein each annotation in the annotation set identifies the target type of the biological structure depicted in the first object or region, and wherein each annotation includes an identifier associated with the annotator who generated the annotation; Identifying a first annotation generated by a first annotator from the plurality of annotations, wherein the first annotation identifies a first target type of the biological structure, and wherein the first annotator is associated with a first weight; Identifying a second annotation generated by a second annotator from the plurality of annotations, wherein the second annotation identifies a second target type of the biological structure, wherein the second annotator is associated with a second weight, and wherein the first target type is different from the second target type; Determining that the first weight is greater than the second weight; And In response to determining that the first weight is greater than the second weight, generating a ground truth label for the first object or region of the image, wherein the ground truth label identifies the first target type of the biological structure.

10. The method according to claim 9, further comprising: Determining a statistical value based on the first weight and the second weight; And Determining whether the statistical value exceeds a confidence threshold.

11. The method according to claim 10, further comprising: In response to determining that the statistical value does not exceed the confidence threshold, generating an output indicating that no consensus has been reached between the first annotator and the second annotator regarding the ground truth label.

12. The method according to claim 10, further comprising: Determining that the statistical value exceeds the confidence threshold; And In response to determining that the statistical value exceeds the confidence threshold, generating an output indicating that a consensus has been reached between the first annotator and the second annotator regarding the ground truth label.

13. The method according to any one of claims 9 to 12, wherein: Receiving an input from a user; and Adjusting the first weight and / or the second weight based on the input.

14. The method according to any one of claims 9 to 13, further comprising: identifying a third annotation generated by a third annotator from the plurality of annotations, wherein the third annotation identifies the second target type of the biological structure, and wherein the third annotator is associated with a third weight; generating another statistical value based on the second weight and the third weight; determining that the another statistical value is greater than the first weight; and in response to determining that the another statistical value is greater than the first weight, redefining the ground truth label for the first object or region of the image to identify the second target type of the biological structure.

15. The method according to any one of claims 9 to 14, further comprising: providing the image with the ground truth label to train the machine learning model, wherein the machine learning model is trained to predict whether another biological structure depicted in another image corresponds to the first target type of the biological structure.

16. The method according to any one of claims 9 to 15, further comprising displaying the image with the ground truth label on a graphical user interface.

17. A system, comprising: one or more data processors; and a non-transitory computer-readable storage medium containing instructions that, when executed on the one or more data processors, cause the one or more data processors to perform any of the steps of the method according to any one of claims 1 to 16.

18. A computer program product tangibly embodied in a non-transitory machine-readable storage medium, the computer program product comprising instructions configured to cause one or more data processors to perform any of the steps of the method according to any one of claims 1 to 16.

Citation Information

Patent Citations

  • Whole slide image registration and cross-image annotation devices, systems and methods

    WO2014140070A2