Expression level prediction for biomarkers in digital pathology images

By using computer methods to process dual immunohistochemical images in digital pathology, the expression level of biomarkers is predicted, and the problem of inaccurate evaluation in the prior art is solved, achieving more accurate and efficient analysis.

CN120035838APending Publication Date: 2025-05-23VENTANA MEDICAL SYSTEMS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380071696.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-10-10
Filing Date
2023-10-05
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

Existing digital pathology analytical methods rely on the subjective scores of pathologists, resulting in inaccurate assessment of biomarker expression levels.

Method used

Using a computer-implemented approach, a synthetic image depicting biomarkers was generated by accessing dual immunohistochemical images of sample slices, and features in the images were processed using a trained machine learning model to predict the expression level of biomarkers.

Benefits of technology

Accurate prediction of biomarker expression levels is achieved, reducing subjectivity and improving analysis efficiency and consistency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure HDA0005348570490000011
    Figure HDA0005348570490000011
  • Figure HDA0005348570490000021
    Figure HDA0005348570490000021
  • Figure HDA0005348570490000031
    Figure HDA0005348570490000031
Patent Text Reader

Abstract

Embodiments disclosed herein generally relate to expression level prediction for digital pathology images. In particular, aspects of the present disclosure relate to acquiring a dual immunohistochemical image of a slice of a sample, where the dual immunohistochemical image includes a cellular depiction associated with a first biomarker and / or a second biomarker corresponding to a disease; generating a first composite image depicting the first biomarker and a second composite image depicting the second biomarker from the dual immunohistochemical image; determining a set of features representative of pixel intensities of the cell depictions in the first composite image and the second composite image; processing the set of features using a trained machine learning model; and outputting a result corresponding to a characterization of the prediction of the sample with respect to the disease based on an output of the processing corresponding to the predicted expression levels of the first biomarker and the second biomarker.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Claiming priority

[0002] This application claims priority to U.S. Provisional Patent Application No. 63 / 414,751, filed on October 10, 2022, entitled “EXPRESSION-LEVEL PREDICTION FOR BIOMARKERS IN DIGITAL PATHOLOGY IMAGES,” and is incorporated by reference in its entirety. Background Art

[0003] Digital pathology involves scanning a slide of a sample (e.g., a tissue sample, a blood sample, a urine sample, etc.) into a digital image. The sample can be stained so that the selected protein (antigen) in the cell is visually differentially labeled relative to the rest of the sample. The target protein in the sample can be referred to as a biomarker. A digital image with one or more stains for a biomarker can be generated for a tissue sample. These digital images can be referred to as a histopathology image. Histopathology images can visualize the spatial relationship between tumor cells and non-tumor cells in a tissue sample. Image analysis can be performed to identify and quantify biomarkers in a tissue sample. Image analysis can be performed by a pathologist to facilitate the characterization of biomarkers (e.g., in terms of expression level, presence, size, shape, and / or position) in order to inform (e.g.) the diagnosis of a disease, the determination of a treatment plan, or the assessment of a response to a therapy. However, the analysis performed by a pathologist may be subjective and inaccurate for scoring the expression level of the biomarkers in the image. Summary of the invention

[0004] Embodiments of the present disclosure relate to techniques for predicting the expression levels of biomarkers in digital pathology images. In some embodiments, a computer-implemented method relates to accessing a dual immunohistochemistry (IHC) image of a sample slice. The dual IHC image includes a cell depiction associated with one or more of a first biomarker and a second biomarker corresponding to a disease. The computer-implemented method further relates to generating a first composite image depicting the first biomarker and a second composite image depicting the second biomarker from the dual IHC image, and determining a set of features for each of the first composite image and the second composite image, the set of features representing the pixel intensity of the cell depiction in the first composite image and the second composite image. The computer-implemented method also relates to processing the set of features using a trained machine learning model. The output of the processing corresponds to the predicted expression level of the first biomarker and the second biomarker. In addition, the computer-implemented method relates to outputting a result corresponding to a characterization of the prediction of the sample about the disease based on the output of the processing.

[0005] In some embodiments, the computer-implemented method further involves, prior to determining the set of features, pre-processing the first composite image and the second composite image by performing color deconvolution on the first composite image and the second composite image.

[0006] In some embodiments, the computer-implemented method further involves, prior to determining the set of features, processing the first composite image and the second composite image using another trained machine learning model. Another output of the processing identifies a first cell depiction of the first composite image predicted to depict the first biomarker and a second cell depiction of the second composite image predicted to depict the second biomarker.

[0007] In some embodiments, determining the set of features for the first composite image involves: determining, for each cell in the first cell depiction, a first metric associated with an intensity value for a block of the cell that includes the cell; aggregating the first metric for each block for the first cell depiction; and determining, based on the aggregation, a plurality of intensity values ​​for the first cell depiction. Each intensity value in the plurality of intensity values ​​corresponds to an intensity percentile, and the plurality of intensity values ​​corresponds to the set of features.

[0008] In some embodiments, determining the set of features for the first composite image involves determining, for each cell in the first cell depiction, a first plurality of intensity values ​​corresponding to intensity percentiles for a block that includes the cell, aggregating the first plurality of intensity values ​​for each block for the first cell depiction to generate a second plurality of intensity values, and determining a set of metrics associated with a distribution of the second plurality of intensity values. The set of metrics corresponds to the set of features.

[0009] In some embodiments, the first biomarker comprises estrogen receptor protein and the second biomarker comprises progesterone receptor protein.

[0010] In some embodiments, the trained machine learning model is a linear regression model.

[0011] In some embodiments, the sample section of the specimen includes a first stain for the first biomarker and a second stain for the second biomarker.

[0012] In some embodiments, the first stain comprises tetramethylrhodamine and the second stain comprises 4-dimethylaminoazobenzene-4'-sulfonyl.

[0013] In some embodiments, the computer-implemented method further involves performing subsequent processing to generate the result of the characterization of the prediction of the sample. Performing the subsequent processing includes detecting a delineation of a group of tumor cells. The result characterizes the presence, number and / or size of the group of tumor cells.

[0014] In some embodiments, a system includes: one or more data processors; and a non-transitory computer-readable storage medium containing instructions that cause the one or more data processors to perform operations when executed on the one or more data processors. These operations include accessing a dual IHC image of a sample slice. The dual IHC image includes a cell depiction associated with one or more of a first biomarker and a second biomarker corresponding to a disease. The operations further include generating a first composite image depicting the first biomarker and a second composite image depicting the second biomarker from the dual IHC image, and determining a set of features for each of the first composite image and the second composite image, the set of features representing the pixel intensity of the cell depiction in the first composite image and the second composite image. The operations also include processing the set of features using a trained machine learning model. The output of the processing corresponds to the predicted expression levels of the first biomarker and the second biomarker. In addition, the operations include outputting a result corresponding to a characterization of the prediction of the sample regarding the disease based on the output of the processing.

[0015] In some embodiments, a computer program product tangibly embodied in a non-transitory machine-readable storage medium includes instructions configured to cause one or more data processors to perform operations. These operations include accessing a dual IHC image of a sample slice. The dual IHC image includes a cell depiction associated with one or more of a first biomarker and a second biomarker corresponding to a disease. These operations further include generating a first composite image depicting the first biomarker and a second composite image depicting the second biomarker from the dual IHC image, and determining a set of features for each of the first composite image and the second composite image, the set of features representing the pixel intensity of the cell depiction in the first composite image and the second composite image. These operations also include processing the set of features using a trained machine learning model. The output of the processing corresponds to the predicted expression levels of the first biomarker and the second biomarker. In addition, these operations include outputting a result corresponding to a characterization of the prediction of the sample regarding the disease based on the output of the processing.

[0016] The terms and expressions that have been adopted are used as terms of description and not of limitation, and when such terms and expressions are used, there is no intention to exclude any equivalents of the features shown and described or portions thereof, but it should be recognized that various modifications are possible within the scope of the claimed invention. Therefore, it should be understood that although the claimed invention has been specifically disclosed through embodiments and optional features, modifications and variations of the concepts disclosed herein may be adopted by those skilled in the art, and such modifications and variations are considered to be within the scope of the invention defined by the appended claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] This patent or application file contains at least one drawing executed in color. Copies of this patent or patent application publication with one or more color drawing(s) will be provided by the Office upon request and payment of the necessary fee.

[0018] Aspects and features of various embodiments will become more apparent by describing examples with reference to the accompanying drawings, in which:

[0019] Figure 1 An exemplary computing system for training and using a machine learning model for expression level prediction is shown;

[0020] Figure 2 Exemplary double immunohistochemistry images of sample sections stained for estrogen receptor protein and progesterone receptor protein are illustrated;

[0021] Figure 3 illustrates the generation of a composite image from an exemplary double immunohistochemistry image of a sample section stained for estrogen receptor protein and progesterone receptor protein;

[0022] Figure 4 An example of an intensity composite image generated by applying color deconvolution to a double immunohistochemistry image is illustrated;

[0023] Figure 5 An example of the output of a trained machine learning model for predicting biomarker profiles is illustrated;

[0024] Figure 6 illustrates an image of cells predicted to depict positive staining for a biomarker superimposed on an intensity composite image for the biomarker;

[0025] Figure 7 An image illustrating predicted biomarker delineation and extracted patches;

[0026] FIG. 8A to FIG. 8B An example table illustrating intensity values ​​of intensity percentiles for each of weakly stained images and medium to highly stained images for estrogen receptor protein;

[0027] Fig. 9 Example histograms corresponding to weakly stained images and medium to highly stained images for estrogen receptor protein are illustrated;

[0028] Fig.10 Example histograms corresponding to weakly stained images and medium to highly stained images for progesterone receptor protein are illustrated;

[0029] FIG. 11A to FIG. 11B Additional example histograms corresponding to weakly stained images and moderately to highly stained images for estrogen receptor protein are illustrated;

[0030] FIG. 12A to FIG. 12B Additional example histograms corresponding to weakly stained images and moderately to highly stained images for progesterone receptor protein are illustrated;

[0031] Fig.13 An exemplary process for expression level prediction of digital pathology images is illustrated;

[0032] FIG. 14A to FIG. 14C Illustrated are example expression levels of Dabsyl estrogen receptor protein determined by three pathologists for 50 fields of view;

[0033] FIG. 15A to FIG. 15C Example expression levels of Tamra progesterone receptor protein determined by three pathologists for 50 fields of view are illustrated;

[0034] FIG. 16A to FIG. 16B illustrates example performance of predicting expression levels using a machine learning model; and

[0035] Fig.17 Exemplary expression level scores generated by a pathologist and a trained machine learning model are illustrated.

[0036] In the drawings, similar parts and / or features may have the same reference number. In addition, various parts of the same type may be distinguished by following the reference number with a dash and a second reference number that distinguishes the similar parts. If only the first reference number is used in the specification, the description applies to any similar part having the same first reference number, regardless of the second reference number. DETAILED DESCRIPTION

[0037] I. Overview

[0038] The present disclosure describes techniques for predicting expression levels of biomarkers in digital pathology images. More specifically, some embodiments of the present disclosure provide for processing dual immunohistochemistry (IHC) images by a machine learning model trained for expression level prediction.

[0039] Digital pathology may involve the interpretation of digitized pathology images in order to correctly diagnose a subject and guide treatment decisions. In a digital pathology solution, an image analysis workflow may be established to automatically detect or classify biological objects of interest (e.g., positive, negative tumor cells, etc.). An exemplary digital pathology solution workflow includes obtaining a tissue slide, scanning a preselected area or the entirety of the tissue slide using a digital image scanner to obtain a digital image, performing image analysis on the digital image using one or more image analysis algorithms, and possibly detecting, quantifying each object of interest (e.g., counting or identifying the object-specific or cumulative area of ​​each object of interest) based on the image analysis (e.g., quantitative or semi-quantitative scores, such as positive, negative, moderate, weak, etc.).

[0040] During imaging and analysis, regions of a digital pathology image can be segmented into target regions (e.g., positive or negative tumor cells) and non-target regions (e.g., normal tissue or blank slide regions). Each target region can include a region of interest that can be characterized and / or quantified. A machine learning model can be developed to segment the target regions. A pathologist can then score the expression levels of the biomarkers in the segmentation. However, pathologist scores can be subjective, and having multiple pathologists score each segment can be time-consuming and resource-intensive.

[0041] In some embodiments, a trained machine learning model may determine a prediction of the expression level of a biomarker in a digital pathology image. A higher expression level may correspond to a higher probability of the presence of a disease. The digital pathology image may be a dual IHC image of a sample slice stained for two biomarkers. A composite image may be generated for each biomarker (e.g., by applying color deconvolution to the dual IHC image). Then, a set of features representing the pixel intensity in the composite image may be determined for each composite image. In some examples, the set of features may be extracted from the composite images, which have been further processed into grayscale images with pixel values ​​representing intensity. For each percentile in a plurality of intensity percentiles, the set of features may be a single intensity value corresponding to the intensity percentile. Alternatively, the set of features may be a set of metrics corresponding to the distribution of intensity values ​​for each intensity percentile. In either case, the trained machine learning model may process the set of features to generate an output corresponding to the predicted expression level of the first biomarker and the second biomarker. Based on the predicted expression level, a characterization of the sample with respect to the disease may be determined. For example, the characterization may be a diagnosis of the disease, a prognosis of the disease, or a predicted therapeutic response to the disease.

[0042] Using features representing pixel intensities as input to a trained machine learning model may produce expression level predictions that accurately correlate with pathologist scores. Thus, a trained machine learning model can provide accurate and fast expression level predictions. Thus, predictions made by a trained machine learning model can result in more efficient and better diagnosis and treatment assessments of diseases (e.g., cancer and / or infectious diseases).

[0043] II. Computing Environment

[0044] Figure 1 An exemplary computing system 100 for training and using a machine learning model for expression level prediction is shown. Images are generated at an image generation system 105. These images can be digital pathology images, such as dual IHC images. The fixing / embedding system 110 uses a fixative (e.g., a liquid fixative, such as a formaldehyde solution) and / or an embedding material (e.g., a histological wax such as paraffin and / or one or more resins such as styrene or polyethylene) to fix and / or embed a tissue sample (e.g., a sample of at least a portion of at least one tumor). Each slice can be fixed by exposing the slice to a fixative for a predetermined period of time (e.g., at least 3 hours) and then dehydrating the slice (e.g., via exposure to an ethanol solution and / or a clarifying intermediate). When the slice is in a liquid state (e.g., when heated), the embedding material can infiltrate the sample.

[0045] The tissue slicer 115 then slices the fixed and / or embedded tissue sample (e.g., tumor sample) to obtain a series of slices, each slice having a thickness of, for example, 4 to 5 microns. Such slicing can be performed by first cooling the sample and slicing the sample in a warm water bath. The tissue can be sliced ​​using, for example, a vibrating microtome or a compression microtome.

[0046] Because tissue sections and cells therein are actually transparent, the preparation of sections usually includes staining (e.g., automatic staining) the tissue sections to make the relevant structures more visible. In some cases, staining is done manually. In some cases, staining is done semi-automatically or automatically using a staining system 120.

[0047] Staining can include exposing a single slice of tissue to one or more different stains (e.g., continuously or simultaneously) to express different features of the tissue. For example, each slice can be exposed to a predefined volume of stain for a predefined time period. Dual determination includes a method in which a slice is stained with two biomarker stains. Single determination includes a method in which a slice is stained with a single biomarker stain. Multiple determination includes a method in which a slice is stained with two or more biomarker stains.

[0048] An exemplary type of tissue staining is a histochemical staining, which uses one or more chemical dyes (e.g., acid dyes, basic dyes) to stain tissue structures. Histochemical staining can be used to indicate general aspects of tissue morphology and / or cell histology (e.g., to distinguish between nuclei and cytoplasm, indicate lipid droplets, etc.). An example of a histochemical stain is hematoxylin and eosin (H&E). Other examples of histochemical stains include trichrome stains (e.g., Masson's trichrome staining), periodic acid-Schiff (PAS), silver stains, and iron stains. The molecular weight of a histochemical staining reagent (e.g., dye) is typically about 500 kilodaltons (kD) or less, although some histochemical staining reagents (e.g., Alcian blue, phosphomolybdic acid (PMA)) may have a molecular weight of up to two or three thousand kD. An example of a high molecular weight histochemical staining reagent is alpha-amylase (about 55kD), which can be used to indicate glycogen.

[0049] Another type of tissue staining is immunohistochemistry (IHC, also called "immunostaining"), which uses a primary antibody that specifically binds to a target antigen of interest (also called a biomarker). IHC can be direct or indirect. In direct IHC, the primary antibody is directly conjugated to a marker (e.g., a chromophore or fluorophore). In indirect IHC, the primary antibody first binds to the target antigen, and then a secondary antibody conjugated to a marker (e.g., a chromophore or fluorophore) binds to the primary antibody. The molecular weight of IHC reagents is much higher than that of histochemical staining reagents, as the molecular weight of antibodies is about 150 kD or higher.

[0050] Then, the segments can be mounted on corresponding slides, respectively, and the imaging system 125 can then scan or image these corresponding slides to generate original digital pathology or histopathology images. Histopathology images can be included in images 130a-n. Each segment can be mounted on a slide, which is then scanned to create a digital image, which can then be examined by digital pathology image analysis and / or interpreted by a human pathologist (e.g., using image viewer software). Pathologists can review and manually annotate the digital images of the slides (e.g., expression levels, tumor areas, necrosis, etc.) to enable the use of image analysis algorithms to extract meaningful quantitative measurements (e.g., to detect and classify the biological object of interest). Traditionally, pathologists can manually annotate each continuous image of multiple tissue sections from a tissue sample to identify the same aspects on each continuous tissue section.

[0051] The computing system 100 may include an analysis system 135 for training and executing a machine learning model. Examples of machine learning models may be a deep convolutional neural network, a U-Net, a V-Net, a residual neural network, a recurrent neural network, a linear regression model, a logistic regression model, or a support vector machine. The machine learning model may be an expression level prediction model 140 that is trained and / or used, for example, to predict the expression level of a biomarker in an image. The expression level of the biomarker may correspond to a diagnosis or treatment decision associated with the disease (e.g., a specific expression level is associated with a predicted positive diagnosis or treatment action). Therefore, the image may be additionally processed based on the predicted expression level to further predict whether the image includes a depiction of a group of tumor cells or other structural and / or functional biological entities associated with the disease, whether the image is associated with a diagnosis of the disease, whether the image is associated with a classification of the disease (e.g., stage, subtype, etc.), and / or the image is associated with a prognosis for the disease. The prediction may characterize the presence, number, and / or size of the group of tumor cells or other structural and / or functional biological entities, the diagnosis of the disease, the classification of the disease, and / or the prognosis of the disease.

[0052] The analysis system 135 may additionally train and execute another machine learning model for predicting the depiction of one or more positively stained biomarkers in an image. Examples of another machine learning model may be a deep convolutional neural network, a U-Net, a V-Net, a residual neural network, a recurrent neural network, a linear regression model, a logistic regression model, or a support vector machine. Another machine learning model may predict positive and negative staining of depictions of biomarkers of cells in an image (e.g., a dual image or a single image). Expression level predictions may be performed only in association with cells with a positive prediction of at least one biomarker, so the output of another machine learning model may be used to determine which portions of the image will be subjected to expression level predictions.

[0053] The training controller 145 may execute code to train the expression level prediction model 140 and / or other machine learning models using one or more training data sets 150. Each training data set 150 may include a set of training images from images 130a-n. Each of these images may include a double IHC image stained for depicting two biomarkers or a single IHC image stained for depicting one of the two biomarkers and one or more biological objects (e.g., a group of cells of one or more types). Each image in the first subset of the set of training images may include one or more biomarkers, and each image in the second subset of the set of training images may not have a biomarker. Each of these images may depict a portion of a sample (such as a tissue sample (e.g., colorectal, bladder, breast, pancreatic, lung or gastric tissue), a blood sample, a urine sample). In some cases, each of one or more of these images depicts a plurality of tumor cells or a plurality of other structural and / or functional biological entities. The training data set 150 may have been collected, for example, from the image generation system 105.

[0054] In some cases, the training controller 145 determines or learns preprocessing parameters and / or methods. For example, preprocessing may include generating a composite image from a dual IHC image, wherein each composite image depicts one of the two biomarkers in the dual IHC image. The dual IHC image may be, for example, an image of a sample slice stained with a first stain (e.g., tetramethylrhodamine (Tamra)) associated with a first biomarker (e.g., progesterone receptor protein) and a second stain (e.g., 4-dimethylaminoazobenzene-4'-sulfonyl (Dabsyl)) associated with a second biomarker (e.g., estrogen receptor protein). In addition, the sample slice may also include a counterstain (e.g., hematoxylin). Color deconvolution may be applied to generate a composite image for each biomarker. That is, a first color vector may be applied to the dual IHC image to generate a first composite image depicting the first biomarker based on the color of the first stain, and a second color vector may be applied to the dual IHC image to generate a second composite image depicting the second biomarker based on the color of the second stain. Figure 2 and Figure 3 An example of a dual IHC image 202 / 302 of a sample section stained for estrogen receptor protein and progesterone receptor protein is illustrated. Figure 3 In FIG. 3 , dual IHC image 302 is a portion of full slide image 301. Color deconvolution is performed on dual IHC images 202 / 302 to generate composite images 204 / 206 / 304 / 306. Composite images 304 / 404 each depict estrogen receptor protein, while composite image 206 / 306 depicts progesterone receptor protein.

[0055] Additionally, each of the composite images may be color deconvolved to generate an image representing intensity (e.g., a grayscale image having pixel values ​​representing intensity between 0 and 255). Color deconvolution may involve determining a stain reference vector from the composite image or the counterstain-free image, performing a matrix inversion using the reference vector to determine the contribution of each stain to the optical density or intensity of the pixel, and generating an intensity composite single image by recombining the unmixed images.

[0056] Figure 4 An example of an intensity composite image generated by applying color deconvolution to a dual IHC image is illustrated. For example, image 412 represents hematoxylin intensity, image 414A represents Dabsyl estrogen receptor intensity, and image 416A represents Tamra progesterone receptor intensity for dual IHC image 402A. In addition, images 414B to 414C represent Dabsyl estrogen receptor intensity, and images 416B to 416C represent Tamra progesterone receptor intensity for dual IHC images 402B to 402C, respectively. Dual IHC image 402B may correspond to medium to high staining, while dual IHC image 402C may correspond to weak staining.

[0057] Back to Figure 1 , the training controller 145 can feed the original image or the pre-processed image (e.g., each of the dual IHC images and / or the composite images) into another trained machine learning model having an architecture (e.g., U-Net) used during previous training and configured with learned parameters. The other trained machine learning model can generate an output that identifies a first cell depiction predicted to depict positive staining of a first biomarker and a second cell depiction predicted to depict positive staining of a second biomarker.

[0058] refer to Figure 5, showing the output of another trained machine learning model. Image 522 illustrates a dual IHC image with predicted biomarker depictions of various colors. For example, cells depicted in red correspond to positive staining for two biomarkers, cells depicted in green correspond to positive staining for a first biomarker (e.g., estrogen receptor protein), cells depicted in blue correspond to positive staining for a second biomarker (e.g., progesterone receptor protein), cells depicted in yellow correspond to negative staining for two biomarkers, and cells depicted in black correspond to other detected cells (e.g., stoma cells). In addition, the composite image can be input into the trained machine learning model to generate image 524 / 526, or the predicted biomarker depiction can be extracted from image 522 and superimposed on the composite image to generate image 524 / 526. In image 524 (which corresponds to a composite image depicting estrogen receptor protein), cells depicted in red correspond to positive staining for estrogen receptor protein, cells depicted in yellow correspond to negative staining for estrogen receptor protein, and cells depicted in black correspond to other detected cells. In image 526 (which corresponds to a composite image depicting progesterone receptor protein), cells depicted in red correspond to positive staining for progesterone receptor protein, cells depicted in yellow correspond to negative staining for progesterone receptor protein, and cells depicted in black correspond to other detected cells.

[0059] Back to Figure 1 Once another trained machine learning model outputs a predicted biomarker depiction, the training controller 145 can generate an input for the expression level prediction model 140. Expression level prediction can be performed only for cells predicted to depict positive staining for at least one of these biomarkers. Therefore, based on the output of another trained machine learning model, a portion of the dual IHC image and / or composite image depicting positive staining for one or more of these biomarkers can be extracted. For example, in an intensity composite image for a first biomarker, a portion predicted to depict positive staining for the first biomarker can be extracted. In addition, in an intensity composite image for a second biomarker, a portion predicted to depict positive staining for a second biomarker can be extracted. Extracting these portions may involve defining a block (e.g., a 5x5 block) around each portion predicted to include positively stained cells.

[0060] Please refer to Figure 6, illustrates an image 632 predicted to depict cells with positive staining for a biomarker superimposed on an intensity composite image for the biomarker. A block 634 is extracted from the image 632. The block 634 is a 5x5 block of a portion of the image 632 predicted to depict cells with positive staining for the biomarker. A plurality of blocks may be extracted from the image 632, and each block may be a 5x5 block surrounding cells predicted to depict positive staining for the biomarker.

[0061] Similarly, Figure 7 An image 722 with predicted biomarker depictions and extracted blocks 730 is illustrated. Intensity composite images 732A to 732C illustrate depictions of cells predicted to include positive staining for biomarkers in image 722, and blocks 734A to 734C illustrate depictions of cells predicted to include positive staining for biomarkers in block 730. Image 732A and block 734A illustrate cells predicted to include depictions of positive staining for both estrogen receptor protein and progesterone receptor protein. Image 732B and block 734B illustrate cells predicted to include depictions of positive staining in the Dabsyl channel, where only cell blocks around cells predicted to include positive staining for estrogen receptor protein are counted, and positively stained progesterone receptor cells are not considered to calculate expression levels. Image 732C and block 734C illustrate depicted cells predicted to include positive staining in the Tamra channel, where only the cell blocks surrounding cells predicted to include positive staining for progesterone receptor protein were counted and positively stained estrogen receptor cells were not considered to calculate expression levels.

[0062] Back to Figure 1In one example, the training controller 145 may perform feature extraction for each block to generate a set of features representing pixel intensities of cell depictions in the first and second composite images. In the first feature extraction technique, the training controller 145 may determine a metric associated with an intensity value for a block including the cell for each cell in the intensity composite image that is predicted to depict positive staining for the first biomarker. For example, the metric may be an average intensity value of pixels in the block. The training controller 145 may then aggregate the metrics for each block in the intensity composite image. Aggregating these metrics may involve sorting the average values ​​for each block from the lowest intensity (e.g., closer to 0) to the highest intensity (e.g., closer to 255). In addition, these metrics may be normalized so that each value is between 0 and 1. From the aggregated metrics, the training controller 145 may determine an intensity value for a cell in the intensity composite image that is predicted to depict positive staining for the first biomarker. Each intensity value may correspond to an intensity percentile from the normalized block intensity. The training controller 145 may perform a similar process for each cell in the intensity composite image predicted to depict positive staining for the second biomarker.

[0063] Go to FIG. 8A to FIG. 8B , illustrates an example table of intensity values ​​of the intensity percentile for each of the weakly stained image 802A and the medium to high stained image 802B for the estrogen receptor protein. Images 804A to 804B are intensity composite images corresponding to the weakly stained image 802A and the medium to high stained image 802B, respectively. The intensity value is the normalized intensity value for the aggregation of each image 804A to 804B. For each of the images 804A to 804B, the normalized intensity value for the aggregation of the 10% intensity percentile, the 25% intensity percentile, the 50% intensity percentile, the 90% intensity percentile, the 95% intensity percentile, the 97.5% intensity percentile, the 99% intensity percentile, the 99.25% intensity percentile, and the 99.5% intensity percentile is illustrated. For image 804A, the aggregated normalized intensity values ​​for the 10% intensity percentile are 0.326535, increasing at each intensity percentile, and 0.763664 for the 99.5% intensity percentile. In contrast, for image 804B, the aggregated normalized intensity values ​​for the 10% intensity percentile are 0.386133, increasing at each intensity percentile, and 0.868026 for the 99.5% intensity percentile. At each intensity percentile, the intensity values ​​are greater for the medium to high stained image 802B than for the weakly stained image 802A.

[0064] Return to Figure 1, an alternative feature extraction technique may involve the training controller 145 determining, for each block in the intensity composite image predicted to depict positive staining for the first biomarker, an intensity value corresponding to an intensity percentile for the block (e.g., 50%, 60%, 70%, 80%, 90%, and 95%). Thus, for a given block, the training controller 145 may determine a distribution of intensity values ​​in the block. Based on the distribution, the training controller 145 may determine an intensity value associated with each intensity percentile. The training controller 145 may then aggregate the intensity values ​​of the intensity percentile for each block in the intensity composite image. That is, the intensity values ​​associated with the 50% percentile for each block may be aggregated, the intensity values ​​associated with the 60% percentile may be aggregated, and so on. The training controller 145 may then calculate a histogram for each intensity percentile and normalize the bins for each histogram.

[0065] Go to Fig. 9 , illustrates examples of histograms corresponding to a weakly stained image 904A and a medium to high stained image 904B for estrogen receptor protein. The weakly stained image 904A is generated from the dual IHC image 902A, and the medium to high stained image 904B is generated from the dual image 902B. For the weakly stained image, most of the intensity values ​​are between 0.3 and 0.7, while for the medium to high stained image, most of the intensity values ​​are between 0.6 and 0.9.

[0066] Fig.10 Example histograms corresponding to a weak staining image 1004A and a medium to high staining image 1004B for progesterone receptor protein are illustrated. The weak staining image 1004A is generated from the dual IHC image 1002A, and the medium to high staining image 1004B is generated from the dual image 1002B. For the weak staining image 1004A, most of the intensity values ​​are between 0.3 and 0.7, while for the medium to high staining image 1004B, most of the intensity values ​​are between 0.6 and 0.9.

[0067] FIG. 11A to FIG. 11B An example histogram corresponding to a weakly stained image 1104A and a medium to high stained image 1104B for estrogen receptor protein is illustrated. The histogram represents the aggregated intensity values ​​for multiple intensity percentiles of the weakly stained image 1104A and the medium to high stained image 1104B. The intensity percentiles include the 50% intensity percentile, the 60% intensity percentile, the 70% intensity percentile, the 80% intensity percentile, the 90% intensity percentile, and the 95% intensity percentile. For each intensity percentile, the intensity value aggregated for the medium to high stained image 1104B is greater than that for the weakly stained image 1104A.

[0068] FIG. 12A to FIG. 12BExample histograms corresponding to weakly stained images 1204A and moderately to highly stained images 1204B for progesterone receptor protein are illustrated. The histograms represent aggregated intensity values ​​for multiple intensity percentiles for the weakly stained images 1204A and the moderately to highly stained images 1204B. The intensity percentiles include the 50% intensity percentile, the 60% intensity percentile, the 70% intensity percentile, the 80% intensity percentile, the 90% intensity percentile, and the 95% intensity percentile. FIG. 11A to FIG. 11B Similarly, for each intensity percentile, the aggregated intensity values ​​are greater for the medium to high stained image 1204B than for the weakly stained image 1204A.

[0069] Back to Figure 1 , the computing system 100 may include a label mapper 160 that maps an image 130 from the imaging system 125 containing a depiction of a biomarker associated with a disease to a label indicating an expression level of the biomarker. The label may be determined based on one or more expression level determinations of the image 130 by a pathologist. For example, one or more pathologists may provide a determination of an H-score corresponding to the expression level of the biomarker in the image, and the H-score may be used as the label. The H-score may be obtained by the following formula: 3x the percentage of strongly stained nuclei + 2x the percentage of moderately stained nuclei + the percentage of weakly stained nuclei. If multiple pathologists provide H-scores, the label may be the mean or median H-score between the pathologists. The label may also include an intensity value determined from feature extraction. For example, for a first feature extraction technique, the label may include an intensity value and a corresponding intensity percentile. For a second feature extraction method, the label may include a distribution of intensity values ​​for each intensity percentile. The mapping data may be stored in a mapping data repository (not shown). The mapping data may identify the expression level mapped to each image.

[0070] In some cases, the labels associated with the training data set 150 may have already been received, or may be derived from data received from the remote system 155. The received data may include, for example, one or more medical records corresponding to a particular subject to which the one or more images 130 correspond. In some cases, the images or scans that are input to the one or more classifier subsystems are received from the remote system 155. For example, the remote system 155 may receive the image 130 from the image generation system 105, and may then transmit the image 130 or scan (e.g., along with a subject identifier and one or more labels) to the analysis system 135.

[0071] The training controller 145 can train the expression level prediction model 140 using the mapping of the training data set 150. More specifically, the training controller 145 can access the architecture of the model, define (fixed) hyperparameters for the model (the hyperparameters are parameters that affect the learning process, such as the learning rate, size / complexity, etc. of the model), and train the model so that a set of parameters are learned. More specifically, the set of parameters can be learned by identifying parameter values ​​associated with low or lowest losses, costs, or errors generated by comparing the predicted output (obtained using a given parameter value) with the actual output. In some cases, the machine learning model can be configured to iteratively fit a new model to improve the estimated accuracy of the output (e.g., a metric or identifier that includes a prediction of the expression level of a biomarker).

[0072] The machine learning (ML) execution handler 165 can use the architecture and learned parameters to process independent data and generate results. For example, the ML execution handler 165 can access dual IHC images that are not represented in the training data set 150. In some embodiments, the generated dual IHC images are stored in a memory device. The imaging system 125 can be used to generate images. In some embodiments, as described herein, the image is generated or obtained from a microscope or other instrument capable of capturing image data of a microscope slide carrying a sample. In some embodiments, the image is generated or obtained using a 2D scanner (such as a scanner capable of scanning an image tile). Alternatively, the image may have been previously generated (e.g., scanned) and stored in a memory device (or, for that matter, retrieved from a server via a communication network).

[0073] In some cases, the dual IHC images may be preprocessed according to a learned or identified preprocessing technique. For example, the ML execution processing program 165 may generate a composite image depicting each of these biomarkers by applying color deconvolution to the dual IHC images. In addition, the ML execution processing program 165 may also generate an intensity composite image by applying additional color deconvolution to each of the composite images. The original image and / or the preprocessed image (e.g., each of the dual IHC images and / or the composite images) may be fed into a trained machine learning model having an architecture (e.g., U-Net) used during training and configured with learned parameters. The trained machine learning model may generate an output that identifies a first cell depiction predicted to depict a first biomarker and a second cell depiction predicted to depict a second biomarker.

[0074] Once the trained machine learning model outputs predicted biomarker depictions, the ML execution processing program 165 can use the architecture and learned parameters of the expression level prediction model 140 to predict the expression level for the biomarker. Expression level prediction can be performed only for cells predicted to depict positive staining for at least one of these biomarkers. Therefore, based on the output of the trained machine learning model, a portion of the dual IHC image and / or composite image depicting positive staining for one or more of these biomarkers can be extracted. For example, in an intensity composite image for a first biomarker, a portion predicted to depict positive staining for the first biomarker can be extracted. In addition, in an intensity composite image for a first biomarker, a portion predicted to depict positive staining for the first biomarker can be extracted. Extracting these portions may involve defining a block (e.g., a 5x5 block) around each portion predicted to include positively stained cells. Then, the ML execution processing program 165 may perform a feature extraction technique on the intensity composite image to determine the intensity value associated with the intensity percentile for each block and for the entire image.

[0075] The original image and / or the pre-processed image (e.g., the dual IHC image, each of the composite images, and / or each of the intensity composite images) and the intensity values ​​can be fed to the expression level prediction model 140, which has an architecture (e.g., a linear regression model) used during training and is configured with learned parameters. The expression level prediction model 140 can generate an output that identifies the predicted expression levels of the first biomarker and the second biomarker.

[0076] In some cases, the image characterizer 170 identifies a characterization of a prediction about a disease for an image based on the execution of image processing. The execution of the expression level prediction model 140 itself may produce a result including a characterization, or the execution may include a result of a characterization that the image characterizer 170 may use to determine a predicted characterization of a sample. For example, the image characterizer 170 may perform subsequent processing that includes characterizing the presence, number, and / or size of a group of tumor cells predicted to be present in the image. The subsequent processing may additionally or alternatively include characterizing the diagnosis of a disease predicted to be present in the image, classifying the disease predicted to be present in the image, and / or predicting the prognosis of the disease predicted to be present in the image. The image characterizer 170 may apply rules and / or transformations to map the predicted expression level and the associated probability and / or confidence to the characterization. As an illustration, if the result includes a probability greater than 50% that the predicted expression level is above a threshold, a first characterization may be assigned, otherwise a second characterization may be assigned.

[0077] The communication interface 175 can collect the results and transmit one or more results (or processed versions thereof) to a user device or other system. For example, the results can be communicated to the remote system 155. In some cases, the communication interface 175 can generate an output that identifies the presence, number, and / or size of the group of tumor cells, the diagnosis of the disease, the classification of the disease, and / or the prognosis of the disease. The output can then be presented and / or transmitted, which can facilitate the display of the output data, such as on a display of a computing device. The results can be used to determine a diagnosis, a treatment plan, or to assess ongoing treatment for tumor cells.

[0078] III. Exemplary Use Cases

[0079] Fig.13 An exemplary process for expression level prediction of digital pathology images is illustrated. The steps of the process can be performed by one or more systems. Other examples can include more steps, fewer steps, different steps, or steps in a different order.

[0080] At block 1305, a dual IHC image of a sample slice is accessed. The dual IHC image may include a depiction of cells associated with one or more of a first biomarker and a second biomarker corresponding to a disease. For example, to identify breast cancer, the first biomarker may be an estrogen receptor protein, and the second biomarker may be a progesterone receptor protein. The sample slice may include a first stain for the first biomarker and a second stain for the second biomarker. As an example, the first stain may be Dabsyl, and the second stain may be Tamra.

[0081] At block 1310, a first composite image and a second composite image are generated. Color deconvolution may be applied to the dual IHC image to generate the first composite image and the second composite image. The first composite image may depict the first biomarker, and the second composite image may depict the second biomarker. Additional pre-processing may also be applied to the composite images. For example, additional color deconvolution may be applied to the first composite image and the second composite image to generate an intensity composite image having grayscale pixels representing the intensity of the cell depictions in the composite images. The composite images may also be input into a trained machine learning model that identifies cell depictions in the first composite image that are predicted to depict the first biomarker and cell depictions in the second composite image that are predicted to depict the second biomarker.

[0082] At block 1315, a set of features representing pixel intensities of cell depictions is determined. Blocks may be generated that each include a depiction of at least one cell predicted to depict a first biomarker or a depiction of at least one cell predicted to depict a second biomarker. For each cell predicted to depict positive staining for the first biomarker, a metric associated with the intensity value for the block including the cell may be determined. For example, the metric may be the average intensity value of the pixels in the block. The metrics for each block in the intensity composite image may then be aggregated and normalized. From the aggregated metrics, an intensity value for a cell in the intensity composite image predicted to depict positive staining for the first biomarker may be determined. Each intensity value may correspond to an intensity percentile from the normalized block intensity. A similar process may be performed for each cell in the intensity composite image predicted to depict positive staining for the second biomarker. An alternative feature extraction technique may involve determining, for each block in an intensity composite image predicted to depict positive staining for a first biomarker, an intensity value corresponding to an intensity percentile (e.g., 50%, 60%, 70%, 80%, 90%, and 95%) for the block. An intensity value associated with each intensity percentile may be determined, and the intensity values ​​of the intensity percentiles for each block in the intensity composite image may be aggregated. A set of metrics associated with the distribution of the aggregated intensity values ​​for the intensity percentiles may be determined. For example, the set of metrics may be determined from a histogram generated for each intensity percentile.

[0083] At block 1320, the set of features is processed using the trained machine learning model. For the first feature extraction technique, the set of features may be intensity values ​​corresponding to different intensity percentiles. For the second feature extraction technique, the set of features may be a set of metrics associated with a distribution of intensity values ​​aggregated for intensity percentiles.

[0084] At block 1325, a result is output corresponding to a predicted characterization of the sample with respect to the disease. For example, the result may be transmitted to another device (e.g., associated with a care provider) and / or displayed. The result may correspond to a predicted characterization of the sample. The result may characterize the presence, number, and / or size of a group of tumor cells in the image, a diagnosis of the disease, a classification of the disease, and / or a prognosis of the disease.

[0085] IV. Exemplary Results

[0086] FIG. 14A to FIG. 14CThe example expression level of Dabsyl estrogen receptor protein determined by three pathologists for 50 visual fields (e.g., dual IHC images) is illustrated. When Dabsyl estrogen receptor protein is scored, there is a higher consistency among the pathologists in the middle to high staining case, and a large difference is found in the low staining case between the three pathologists. The expression level determined by each pathologist in the three pathologists is compared with the median estrogen receptor protein expression level among the pathologist for each visual field. As shown in the figure, there is a high consistency among the three pathologists, and wherein each pathology has a correlation coefficient between 0.93 and 0.98.

[0087] FIG. 15A to FIG. 15C The example expression level of Tamra progesterone receptor protein determined by three pathologists for 50 visual fields (e.g., dual IHC images) is illustrated. When the Tamra progesterone receptor protein is scored, the medium to high staining cases have a higher consistency among the pathologists, while the low staining cases have a larger difference between the three pathologists. The expression level determined by each pathologist in the three pathologists is compared with the median progesterone receptor protein expression level among the pathologists for each visual field. As shown in the figure, there is a high consistency among the three pathologists, wherein each pathology has a correlation coefficient between 0.88 and 0.96.

[0088] FIG. 16A to FIG. 16B Example performance of using a machine learning model to predict expression levels is illustrated. The intensity values ​​are extracted using the second feature extraction technique described herein above. The trained machine learning model (e.g., Figure 1The expression level prediction model 140 in (140) achieves a high consistency about the expression level determined by the pathologist when predicting the expression level. Compared with the median consensus expression level, the trained machine learning model predicts the expression level, and achieves a higher consistency than the score performed by the pathologist. Compared with the median consensus expression level, the correlation for predicting the estrogen receptor protein expression level by the trained machine learning model is 0.9788, and compared with the median consensus expression level, the correlation for predicting the progesterone receptor protein expression level by the trained machine learning model is 0.9292. The determination of the predicted R square is 0.958 and 0.8635 respectively. The table also shows the correlation between the three pathologists and the trained machine learning model. When the Dabysl estrogen receptor protein expression level is scored, the trained machine learning model achieves a higher correlation with the median consensus than any one of the pathologists. In addition, when the Tamra progesterone receptor protein expression level is scored, the trained machine learning model is superior to two pathologists among the pathologists in terms of correlation with the median consensus. Therefore, the trained machine learning models can produce more consistently accurate predictions of expression levels than pathologists, facilitating more accurate characterization of the disease.

[0089] Fig.17 Exemplary expression level scores generated by a pathologist and a trained machine learning model are illustrated. For image 1702, the pathologist determined that the expression level for estrogen receptor protein was 2.15, while the trained machine learning model determined that the expression level for estrogen receptor protein was 2.10. For image 1704, the pathologist determined that the expression level for estrogen receptor protein was 1.15, and the trained machine learning model determined that the expression level for estrogen receptor protein was 1.16. For image 1706, the pathologist determined that the expression level for progesterone receptor protein was 2.40, while the trained machine learning model determined that the expression level for progesterone receptor protein was 2.35. For image 1708, the pathologist determined that the expression level for progesterone receptor protein was 1.50, while the trained machine learning model determined that the expression level for progesterone receptor protein was 1.53. In each case, the predictions made by the trained machine learning model were within 0.05 of the scores determined by the pathologists, further illustrating the accuracy of the trained machine learning model in predicting the expression levels of biomarkers in digital pathology images.

[0090] V. Other considerations

[0091] Some embodiments of the present disclosure include a system including one or more data processors. In some embodiments, the system includes a non-transitory computer-readable storage medium containing instructions, which when executed on the one or more data processors, causes the one or more data processors to perform part or all of one or more methods disclosed herein and / or part or all of one or more processes disclosed herein. Some embodiments of the present disclosure include a computer program product tangibly embodied in a non-transitory machine-readable storage medium, which includes instructions configured to cause one or more data processors to perform part or all of one or more methods disclosed herein and / or part or all of one or more processes disclosed herein.

[0092] The terms and expressions that have been adopted are used as terms of description rather than limitation, and when such terms and expressions are used, there is no intention to exclude any equivalents of the features shown and described or parts thereof, but it should be recognized that various modifications are possible within the scope of the invention claimed. Therefore, it should be understood that although the claimed invention is specifically disclosed through embodiments and optional features, those skilled in the art may make modifications and changes to the concepts disclosed herein, and such modifications and changes are considered to be within the scope of the invention defined by the appended claims.

[0093] This description only provides preferred exemplary embodiments and is not intended to limit the scope, applicability or configuration of the present disclosure. On the contrary, the description of the preferred exemplary embodiments will provide a feasible description for implementing various embodiments for those skilled in the art. It should be understood that various changes can be made to the function and arrangement of the elements without departing from the spirit and scope set forth in the appended claims.

[0094] Specific details are given in the following description to provide a thorough understanding of the embodiments. However, it should be understood that the embodiments can be practiced without these specific details. For example, circuits, systems, networks, processes, and other components can be shown as parts in block diagram form to avoid obfuscating the embodiments in unnecessary details. In other cases, well-known circuits, processes, algorithms, structures, and techniques can be shown without unnecessary details to avoid obfuscating the embodiments.

Claims

1. A computer-implemented method, include: acquiring a dual immunohistochemistry (IHC) image of the sample section, wherein the dual IHC image includes cell depictions associated with one or more of a first biomarker and a second biomarker corresponding to a disease; generating a first composite image depicting the first biomarker and a second composite image depicting the second biomarker from the dual IHC image; determining, for each of the first composite image and the second composite image, a set of features representing pixel intensities of the cell depictions in the first composite image and the second composite image; processing the set of features using a trained machine learning model, wherein an output of the processing corresponds to a predicted expression level of the first biomarker and the second biomarker; and Based on the output of the processing, a result corresponding to a characterization of a prediction of the sample with respect to the disease is output.

2. The computer-implemented method of claim 1 , further comprising, before determining the set of features: The first composite image and the second composite image are pre-processed by applying color deconvolution to the first composite image and the second composite image.

3. The computer-implemented method of claim 1 , further comprising, before determining the set of features: The first and second composite images are processed using another trained machine learning model, wherein another output of the processing identifies a first cell depiction in the first composite image predicted to depict the first biomarker and a second cell depiction in the second composite image predicted to depict the second biomarker.

4. The computer-implemented method of claim 3, wherein the set of features is determined for the first composite image include: determining, for each cell in the first cell delineation, a first metric associated with an intensity value for a block of the cells including the cell; aggregating the first metric for each block for the first cell depiction; as well as A plurality of intensity values ​​depicted for the first cell is determined based on the aggregation, wherein each intensity value in the plurality of intensity values ​​corresponds to an intensity percentile, and wherein the plurality of intensity values ​​corresponds to the set of features.

5. The computer-implemented method of claim 3, wherein the set of features is determined for the first composite image include: determining, for each cell in the first cell delineation, a first plurality of intensity values ​​corresponding to intensity percentiles for a block including the cell; aggregating the first plurality of intensity values ​​for each block for the first cell depiction to generate a second plurality of intensity values; as well as A set of metrics associated with the distribution of the second plurality of intensity values ​​is determined, wherein the set of metrics corresponds to the set of features. 6 . The computer-implemented method of claim 1 , wherein the first biomarker comprises estrogen receptor protein and the second biomarker comprises progesterone receptor protein.

7. The method of claim 1, wherein the trained machine learning model comprises a linear regression model.

8. The computer-implemented method of claim 1, wherein a sample section of the specimen comprises a first stain for the first biomarker and a second stain for the second biomarker.

9. The computer-implemented method of claim 8, wherein the first stain comprises tetramethylrhodamine and the second stain comprises 4-dimethylaminoazobenzene-4'-sulfonyl.

10. The computer-implemented method of claim 1 , further comprising performing subsequent processing to generate the result of the predicted characterization of the sample, wherein performing the subsequent processing comprises detecting a delineation of a group of tumor cells, and wherein the result characterizes the presence, number and / or size of the group of tumor cells.

11. A system, wherein include: one or more data processors; as well as A non-transitory computer-readable storage medium comprising instructions that, when executed on the one or more data processors, cause the one or more data processors to perform operations including: acquiring a dual immunohistochemistry (IHC) image of the sample section, wherein the dual IHC image includes cell depictions associated with one or more of a first biomarker and a second biomarker corresponding to a disease; generating a first composite image depicting the first biomarker and a second composite image depicting the second biomarker from the dual IHC image; determining, for each of the first composite image and the second composite image, a set of features representing pixel intensities of the cell depictions in the first composite image and the second composite image; processing the set of features using a trained machine learning model, wherein an output of the processing corresponds to a predicted expression level of the first biomarker and the second biomarker; and Based on the output of the processing, a result corresponding to a characterization of a prediction of the sample with respect to the disease is output.

12. The system of claim 11, wherein the non-transitory computer readable medium further comprises instructions that, when executed on the one or more data processors, cause the one or more data processors to perform operations comprising, before determining the set of features: The first composite image and the second composite image are pre-processed by applying color deconvolution to the first composite image and the second composite image.

13. The system of claim 11, wherein the non-transitory computer readable medium further comprises instructions that, when executed on the one or more data processors, cause the one or more data processors to perform operations comprising, before determining the set of features: The first and second composite images are processed using another trained machine learning model, wherein another output of the processing identifies a first cell depiction in the first composite image predicted to depict the first biomarker and a second cell depiction in the second composite image predicted to depict the second biomarker.

14. The system of claim 13, wherein the set of features is determined for the first composite image include: determining, for each cell in the first cell delineation, a first metric associated with an intensity value for a block of the cells including the cell; aggregating the first metric for each block for the first cell depiction; as well as A plurality of intensity values ​​depicted for the first cell is determined based on the aggregation, wherein each intensity value in the plurality of intensity values ​​corresponds to an intensity percentile, and wherein the plurality of intensity values ​​corresponds to the set of features.

15. The system of claim 13, wherein the set of features is determined for the first composite image include: determining, for each cell in the first cell delineation, a first plurality of intensity values ​​corresponding to intensity percentiles for a block including the cell; aggregating the first plurality of intensity values ​​for each block for the first cell depiction to generate a second plurality of intensity values; as well as A set of metrics associated with the distribution of the second plurality of intensity values ​​is determined, wherein the set of metrics corresponds to the set of features.

16. The system of claim 13, wherein the first biomarker comprises estrogen receptor protein and the second biomarker comprises progesterone receptor protein.

17. The system of claim 11, wherein the trained machine learning model comprises a linear regression model.

18. The system of claim 11, wherein the sample section of the specimen comprises a first stain for the first biomarker and a second stain for the second biomarker.

19. The system of claim 18, wherein the first stain comprises tetramethylrhodamine and the second stain comprises 4-dimethylaminoazobenzene-4'-sulfonyl.

20. The system of claim 11, wherein the non-transitory computer readable medium further comprises instructions that, when executed on the one or more data processors, cause the one or more data processors to perform operations comprising: Subsequent processing is performed to generate the result of the predicted characterization of the sample, wherein the subsequent processing includes detecting a delineation of a group of tumor cells, and wherein the result characterizes the presence, number and / or size of the group of tumor cells.

21. A computer program product tangibly embodied in a non-transitory machine-readable storage medium, comprising instructions configured to cause one or more data processors to perform operations comprising: acquiring a dual immunohistochemistry (IHC) image of the sample section, wherein the dual IHC image includes cell depictions associated with one or more of a first biomarker and a second biomarker corresponding to a disease; generating a first composite image depicting the first biomarker and a second composite image depicting the second biomarker from the dual IHC image; determining, for each of the first composite image and the second composite image, a set of features representing pixel intensities of the cell depictions in the first composite image and the second composite image; processing the set of features using a trained machine learning model, wherein an output of the processing corresponds to a predicted expression level of the first biomarker and the second biomarker; and Based on the output of the processing, a result corresponding to a characterization of a prediction of the sample with respect to the disease is output.