Prediction of biomarker expression levels in digital pathology images

A computer-implemented method using machine learning models on dual immunohistochemistry images accurately predicts biomarker expression levels, addressing the subjectivity and inefficiency of pathologist assessments, improving disease diagnosis and treatment.

JP2025533848APending Publication Date: 2025-10-09VENTANA MEDICAL SYSTEMS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025519717
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-10-10
Filing Date
2023-10-05
Publication Date
2025-10-09

AI Technical Summary

Technical Problem

Analysis of biomarker expression levels in digital pathology images by pathologists is subjective and inaccurate, leading to time- and resource-intensive assessments.

Method used

A computer-implemented method using dual immunohistochemistry images and machine learning models to generate composite images and determine pixel intensity features, predicting biomarker expression levels accurately.

Benefits of technology

Facilitates accurate and efficient prediction of biomarker expression levels, enhancing disease diagnosis and treatment evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025533848000001
    Figure 2025533848000001
  • Figure 2025533848000002
    Figure 2025533848000002
  • Figure 2025533848000003
    Figure 2025533848000003
Patent Text Reader

Abstract

[0003] Embodiments disclosed herein generally relate to predicting expression levels of digital pathology images. In particular, aspects of the present disclosure are directed to: accessing a dual immunohistochemistry image of a slice of a specimen, the dual immunohistochemistry image including cellular depictions associated with a first biomarker and / or a second biomarker corresponding to a disease; generating a first composite image depicting the first biomarker and a second composite image depicting the second biomarker from the dual immunohistochemistry image; determining a set of features representing pixel intensities of cellular depictions in the first composite image and the second composite image; processing the set of features using a trained machine learning model; and outputting a result corresponding to a predicted characterization of the specimen for disease based on the output of the processing corresponding to a predicted expression level of the first biomarker and the second biomarker.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Priority claims This application claims priority to U.S. Provisional Patent Application No. 63 / 414,751, filed October 10, 2022, entitled "EXPRESSION-LEVEL PREDICTION FOR BIOMARKERS IN DIGITAL PATHOLOGY IMAGES," which is incorporated by reference in its entirety. [Background technology]

[0002] Digital pathology involves scanning slides of samples (e.g., tissue samples, blood samples, urine samples, etc.) into digital images. The samples can be stained so that selected proteins (antigens) in the cells are visually marked differentially from the rest of the sample. The target proteins of the specimen can be called biomarkers. Digital images with one or more stains for the biomarkers can be generated for the tissue sample. These digital images are sometimes called histopathological images. The histopathological images can enable visualization of the spatial relationship between tumor cells and non-tumor cells in the tissue sample. Image analysis can be performed to identify and quantify biomarkers in the tissue sample. Image analysis can be performed by a pathologist to facilitate characterization of biomarkers (e.g., with respect to expression level, presence, size, shape, and / or location) to inform (for example) disease diagnosis, determination of treatment plan, or evaluation of response to therapy. However, analysis performed by a pathologist can be subjective and inaccurate in scoring the expression levels of biomarkers in the images. Summary of the Invention

[0003]

[0003] Embodiments of the present disclosure relate to techniques for predicting expression levels of biomarkers in digital pathology images. In some embodiments, a computer-implemented method includes accessing dual immunohistochemistry (IHC) images of slices of a specimen. The dual IHC images include cellular depictions associated with one or more first and second biomarkers corresponding to a disease. The computer-implemented method further includes generating a first composite image depicting the first biomarker and a second composite image depicting the second biomarker from the dual IHC images, and determining, for each of the first and second composite images, a set of features representing pixel intensities of cellular depictions in the first and second composite images. The computer-implemented method also includes processing the set of features using a trained machine learning model. Output of the processing corresponds to predicted expression levels of the first and second biomarkers. The computer-implemented method further includes outputting a result corresponding to a predicted characterization of the specimen with respect to the disease based on the output of the processing.

[0004] In some embodiments, the computer-implemented method further includes preprocessing the first composite image and the second composite image by applying color deconvolution to the first composite image and the second composite image before determining the set of features.

[0005] In some embodiments, the computer-implemented method further includes processing the first composite image and the second composite image using another trained machine learning model before determining the set of features, wherein another output of the processing identifies a first cell representation in the first composite image that is predicted to depict a first biomarker and a second cell representation in the second composite image that is predicted to depict a second biomarker.

[0006] In some embodiments, determining the set of features of the first composite image includes determining, for each cell in the first cell representation, a first metric associated with intensity values ​​of cellular patches that include the cell, aggregating the first metrics for each patch for the first cell representation, and determining a plurality of intensity values ​​for the first cell representation based on the aggregating, wherein each intensity value of the plurality of intensity values ​​corresponds to an intensity percentile, and the plurality of intensity values ​​correspond to the set of features.

[0007] In some embodiments, determining the set of features of the first composite image includes determining, for each cell of the first cell representation, a first plurality of intensity values ​​corresponding to intensity percentiles of patches containing the cell, aggregating the first plurality of intensity values ​​of each patch for the first cell representation to generate a second plurality of intensity values, and determining a set of metrics associated with a distribution of the second plurality of intensity values, wherein the set of metrics corresponds to the set of features.

[0008] In some embodiments, the first biomarker comprises an estrogen receptor protein and the second biomarker comprises a progesterone receptor protein.

[0009] In some embodiments, the trained machine learning model is a linear regression model.

[0010] In some embodiments, a sample slice of the specimen comprises a first stain for a first biomarker and a second stain for a second biomarker.

[0011] In some embodiments, the first stain comprises tetramethylrhodamine and the second stain comprises 4-dimethylaminoazobenzene-4'-sulfonyl.

[0012] In some embodiments, the computer-implemented method further includes performing subsequent processing to generate predicted characterization results for the specimen. Performing the subsequent processing includes detecting a set of tumor cell delineations. The results characterize the presence, amount, and / or size of the set of tumor cells.

[0013] In some embodiments, the system includes one or more data processors and a non-transitory computer-readable storage medium including instructions, which, when executed on the one or more data processors, cause the one or more data processors to perform operations. The operations include accessing a dual IHC image of a slice of a specimen. The dual IHC image includes cellular depictions associated with one or more first and second biomarkers corresponding to a disease. The operations further include generating a first composite image depicting the first biomarker and a second composite image depicting the second biomarker from the dual IHC image, and determining, for each of the first and second composite images, a set of features representing pixel intensities of cellular depictions in the first and second composite images. The operations also include processing the set of features using a trained machine learning model. Output of the processing corresponds to predicted expression levels of the first and second biomarkers. Additionally, the operations include outputting a result corresponding to a predicted characterization of the specimen with respect to the disease based on the output of the processing.

[0014] In some embodiments, a computer program product tangibly embodied in a non-transitory machine-readable storage medium includes instructions configured to cause one or more data processors to perform operations. The operations include accessing a dual IHC image of a slice of a specimen. The dual IHC image includes cellular depictions associated with one or more first and second biomarkers corresponding to a disease. The operations further include generating a first composite image depicting the first biomarker and a second composite image depicting the second biomarker from the dual IHC image, and determining, for each of the first and second composite images, a set of features representing pixel intensities of cellular depictions in the first and second composite images. The operations also include processing the set of features using a trained machine learning model. Output of the processing corresponds to predicted expression levels of the first and second biomarkers. Additionally, the operations include outputting a result corresponding to a predicted characterization of the specimen with respect to the disease based on the output of the processing.

[0015] The terms and expressions which have been employed are used as terms of description rather than of limitation, and there is no intention in the use of such terms and expressions to exclude equivalents of the features shown and described, or portions thereof, but it is recognized that various modifications are possible within the scope of the invention as claimed. Thus, although the claimed invention has been specifically disclosed by embodiments and optional features, it will be understood that modifications and variations of the concepts disclosed herein may be employed by those skilled in the art, and that such modifications and variations are deemed to be within the scope of the invention as defined by the appended claims. [Brief explanation of the drawings]

[0016] The patent or application file contains at least one drawing executed in color. Copies of this patent or patent application publication with color drawing(s) will be provided by the Office upon request and payment of the necessary fee.

[0017] Aspects and features of various embodiments will become more apparent by way of example only and with reference to the accompanying drawings.

[0018] [Figure 1] 1 illustrates an exemplary computing system for training and using machine learning models for expression level prediction.

[0019] [Figure 2] 1 shows an exemplary double immunohistochemistry image of a slice of a specimen stained for estrogen receptor protein and progesterone receptor protein.

[0020] [Figure 3] 1 shows a composite image generated from dual immunohistochemistry images of slices of a specimen stained for estrogen receptor protein and progesterone receptor protein.

[0021] [Figure 4] 1 shows an example of an intensity composite image generated by applying color deconvolution to a dual immunohistochemistry image.

[0022] [Figure 5] 1 shows an example of the output of a trained machine learning model for predicting biomarker portrayals.

[0023] [Figure 6] 1 shows an image of a cell predicted to exhibit positive staining for a biomarker overlaid on an intensity composite image of the biomarker.

[0024] [Figure 7] 1 shows predicted biomarker depictions and images of extracted patches.

[0025] [Figure 8A]1 shows an exemplary table of intensity percentile intensity values ​​for weakly stained and medium-highly stained images of estrogen receptor protein, respectively. [Figure 8B] 1 shows an exemplary table of intensity percentile intensity values ​​for weakly stained and medium-highly stained images of estrogen receptor protein, respectively.

[0026] [Figure 9] 1 shows exemplary histograms corresponding to weakly and moderately stained images of estrogen receptor protein.

[0027] [Figure 10] 1 shows exemplary histograms corresponding to weakly and moderately stained images of progesterone receptor protein.

[0028] [Figure 11A] Additional exemplary histograms corresponding to weakly and moderately stained images of estrogen receptor protein are shown. [Figure 11B] Additional exemplary histograms corresponding to weakly and moderately stained images of estrogen receptor protein are shown.

[0029] [Figure 12A] 10 shows additional exemplary histograms corresponding to weakly and moderately stained images of progesterone receptor protein. [Figure 12B] 10 shows additional exemplary histograms corresponding to weakly and moderately stained images of progesterone receptor protein.

[0030] [Figure 13] 1 illustrates an exemplary process for expression level prediction of digital pathology images.

[0031] [Figure 14A] Exemplary expression levels of Dabcyl estrogen receptor protein as determined by three pathologists for 50 fields are shown. [Figure 14B] Exemplary expression levels of Dabcyl estrogen receptor protein as determined by three pathologists for 50 fields are shown. [Figure 14C] Exemplary expression levels of Dabcyl estrogen receptor protein as determined by three pathologists for 50 fields are shown.

[0032] [Figure 15A] Exemplary expression levels of Tamra progesterone receptor protein as determined by three pathologists for 50 fields are shown. [Figure 15B] Exemplary expression levels of Tamra progesterone receptor protein as determined by three pathologists for 50 fields are shown. [Figure 15C] Exemplary expression levels of Tamra progesterone receptor protein as determined by three pathologists for 50 fields are shown.

[0033] [Figure 16A] 1 shows exemplary performance of predicting expression levels using a machine learning model. [Figure 16B] 1 shows exemplary performance of predicting expression levels using a machine learning model.

[0034] [Figure 17] 1 shows exemplary expression level scores generated by a pathologist and a trained machine learning model.

[0035] In the accompanying drawings, similar components and / or features may have the same reference label. Furthermore, various components of the same type may be distinguished by following the reference label with a dash and a second label that distinguishes between the similar components. If only a first reference label is used herein, the description is applicable to any one of the similar components having the same first reference label, regardless of the second reference label. DETAILED DESCRIPTION OF THE INVENTION

[0036] I. Overview The present disclosure describes techniques for predicting expression levels of biomarkers in digital pathology images. More specifically, some embodiments of the present disclosure provide for processing dual immunohistochemistry (IHC) images with a machine learning model trained for expression level prediction.

[0037] Digital pathology can involve the interpretation of digitized pathology images to accurately diagnose subjects and guide treatment decision-making. In a digital pathology solution, an image analysis workflow can be established to automatically detect or classify biological objects of interest, such as positive or negative tumor cells. An exemplary digital pathology solution workflow includes acquiring a tissue slide, scanning a preselected area or the entire tissue slide with a digital image scanner to obtain a digital image, performing image analysis on the digital image using one or more image analysis algorithms, and potentially detecting and quantifying each object of interest (e.g., counting or identifying its specific or cumulative area) based on the image analysis (e.g., quantitative or semi-quantitative scoring such as positive, negative, moderate, weak, etc.).

[0038] During imaging and analysis, regions of the digital pathology image can be segmented into target regions (e.g., positive and negative tumor cells) and non-target regions (e.g., normal tissue or blank slide regions). Each target region can contain a region of interest that can be characterized and / or quantified. Machine learning models can be developed to segment the target regions. A pathologist can then score the expression levels of biomarkers in the segments. However, pathologist scores can be subjective, and having multiple pathologists score each segment can be time- and resource-intensive.

[0039] In some embodiments, the trained machine learning model determines predicted expression levels of biomarkers in digital pathology images. A higher expression level may correspond to a higher likelihood of disease presence. The digital pathology images may be dual IHC images of specimen slices stained for two biomarkers. A composite image may be generated for each biomarker (e.g., by applying color deconvolution to the dual IHC images). A set of features representing pixel intensities in the composite image may then be determined for each composite image. In some examples, the set of features may be extracted from a composite image that has been further processed into a grayscale image with pixel values ​​representing intensity. The set of features may be a single intensity value corresponding to the intensity percentile for each of multiple intensity percentiles. Alternatively, the set of features may be a set of metrics corresponding to the distribution of intensity values ​​for each intensity percentile. In either case, the trained machine learning model may process the set of features to generate an output corresponding to predicted expression levels of the first biomarker and the second biomarker. Based on the predicted expression levels, a characterization of the specimen for disease may be determined. For example, the characterization can be a diagnosis of a disease, a prognosis of a disease, or a predicted treatment response of a disease.

[0040] Using the feature representing pixel intensity as input to trained machine learning model can lead to expression level prediction that accurately correlates with pathologist scoring.Therefore, trained machine learning model can lead to accurate and faster expression level prediction.Therefore, the prediction made by trained machine learning model can lead to more efficient and better diagnosis and treatment evaluation of disease (for example, cancer and / or infectious disease).

[0041] II. Computing Environment FIG. 1 illustrates an exemplary computing system 100 for training and using a machine learning model for expression level prediction. Images are generated by an image generation system 105. The images may be digital pathology images, such as dual IHC images. A fixation / embedding system 110 fixes and / or embeds a tissue sample (e.g., a sample containing at least a portion of at least one tumor) using a fixative (e.g., a liquid fixative such as a formaldehyde solution) and / or an embedding substance (e.g., a histology wax such as paraffin wax, and / or one or more resins such as styrene or polyethylene). Each slice may be fixed by exposing the slice to a fixative for a predetermined period of time (e.g., at least 3 hours) and then dehydrating the slice (e.g., via exposure to an ethanol solution and / or a clearing intermediate). The embedding substance can penetrate the slice when it is in a liquid state (e.g., when heated).

[0042] A tissue slicer 115 then slices the fixed and / or embedded tissue sample (e.g., a tumor sample) to obtain a series of sections, each having a thickness of, for example, 4-5 microns. Such sectioning may be performed by first chilling the sample and then slicing the sample in a warm water bath. The tissue may be sliced ​​using (for example) a vibratome or compressome.

[0043] Because tissue sections and the cells therein are largely transparent, slide preparation typically involves staining (e.g., automatically staining) the tissue sections to make relevant structures more visible. In some cases, staining is performed manually. In other cases, staining is performed semi-automatically or automatically using staining system 120.

[0044] Staining can involve exposing individual sections of tissue to one or more different stains (e.g., sequentially or simultaneously) to reveal different characteristics of the tissue. For example, each section can be exposed to a predetermined amount of stain for a predetermined period of time. A dual assay involves staining a slide with two biomarker stains. A single assay involves staining a slide with a single biomarker stain. A multiplex assay involves staining a slide with two or more biomarker stains.

[0045] One typical type of tissue stain is histochemical staining, which uses one or more chemical dyes (e.g., acidic dyes, basic dyes) to stain tissue structures. Histochemical stains can be used to reveal general aspects of tissue morphology and / or cellular microanatomy (e.g., distinguishing cell nuclei from cytoplasm, revealing lipid droplets, etc.). One example of a histochemical stain is hematoxylin and eosin (H&E). Other examples of histochemical stains include trichrome stains (e.g., Masson's trichrome), periodic acid-Schiff (PAS), silver stains, and iron stains. The molecular weight of histochemical stains (e.g., dyes) is generally about 500 kilodaltons (kD) or less, although some histochemical stains (e.g., Alcian blue, phosphomolybdic acid (PMA)) can have molecular weights of up to 2000 or 3000 kD. An example of a high molecular weight histochemical stain is α-amylase (approximately 55 kD) which may be used to demonstrate glycogen.

[0046] Another type of tissue staining is immunohistochemistry (IHC, also called "immunostaining"), which uses a primary antibody that specifically binds to a target antigen of interest (also called a biomarker). IHC can be direct or indirect. In direct IHC, the primary antibody is directly conjugated to a label (e.g., a chromophore or fluorophore). In indirect IHC, the primary antibody is first bound to the target antigen, and then a secondary antibody conjugated to a label (e.g., a chromophore or fluorophore) is bound to the primary antibody. Because antibodies have a molecular weight of approximately 150 kD or greater, the molecular weight of IHC reagents is much larger than that of histochemical staining reagents.

[0047] The sections may then be individually mounted onto corresponding slides, and imaging system 125 may then scan or image the slides to generate raw digital pathology, or histopathology, images. The histopathology images may be included in images 130a-n. Each section may be mounted on a slide, which may then be scanned to create a digital image, which may then be examined by digital pathology image analysis and / or by a human pathologist (e.g., using image viewer software). The pathologist may review and manually annotate the digital images of the slides (e.g., expression levels, tumor area, necrosis, etc.) to enable the use of image analysis algorithms to extract meaningful quantitative measures (e.g., to detect and classify biological objects of interest). Traditionally, a pathologist may manually annotate each sequential image of multiple tissue sections from a tissue sample to identify the same aspect for each sequential tissue section.

[0048] The computing system 100 may include an analysis system 135 for training and executing machine learning models. Examples of machine learning models may be deep convolutional neural networks, U-nets, V-nets, residual neural networks, recurrent neural networks, linear regression models, logistic regression models, or support vector machines. The machine learning model may be (for example) an expression level prediction model 140 that is trained and / or used to predict expression levels of biomarkers in an image. The expression levels of the biomarkers may correspond to diagnostic or treatment decisions related to a disease (e.g., a particular expression level is associated with a predicted positive diagnosis or treatment action). As such, additional processing can be performed on the images based on the predicted expression levels to further predict whether the image includes depictions of a set of tumor cells or other structural and / or functional biological entities associated with the disease, whether the image is associated with a diagnosis of the disease, whether the image is associated with a classification of the disease (e.g., stage, subtype, etc.), and / or whether the image is associated with a prognosis of the disease. The prediction may characterize the presence, amount, and / or size of a set of tumor cells or other structural and / or functional biological entities, disease diagnosis, disease classification, and / or disease prognosis.

[0049] The analysis system 135 may additionally train and run another machine learning model to predict the delineation of one or more positive staining biomarkers in the image. Examples of other machine learning models may be deep convolutional neural networks, U-nets, V-nets, residual neural networks, recurrent neural networks, linear regression models, logistic regression models, or support vector machines. The other machine learning model may predict positive and negative staining delineating biomarkers for cells in an image (e.g., a double image or a single image). Because expression level prediction may only be performed with respect to cells with a positive prediction for at least one biomarker, the output of the other machine learning model can be used to determine which portions of the image to perform expression level prediction on.

[0050] The training controller 145 can execute code for training the expression level prediction model 140 and / or other machine learning models using one or more training datasets 150. Each training dataset 150 can include a set of training images from the images 130a-n. Each image can include a double IHC image stained to depict two biomarkers, or a single IHC image stained to depict two biomarkers and one of one or more biological objects (e.g., a set of one or more types of cells). Each image in a first subset of the set of training images can include one or more biomarkers, and each image in a second subset of the set of training images can lack a biomarker. Each of the images can depict a portion of a sample, such as a tissue sample (e.g., colorectal tissue, bladder tissue, breast tissue, pancreatic tissue, lung tissue, or stomach tissue), a blood sample, or a urine sample. In some cases, each of the one or more images depicts multiple tumor cells or multiple other structural and / or functional biological entities. The training data set 150 may have been collected (for example) from an image generation system 105 .

[0051] In some cases, the training controller 145 determines or learns preprocessing parameters and / or approaches. For example, preprocessing can include generating composite images from dual IHC images, where each composite image depicts one of the two biomarkers in the dual IHC images. The dual IHC images can be (for example) images of specimen slices stained with a first stain (e.g., tetramethylrhodamine (Tamra)) associated with a first biomarker (e.g., progesterone receptor protein) and a second stain (e.g., 4-dimethylaminoazobenzene-4'-sulfonyl (dabcyl)) associated with a second biomarker (e.g., estrogen receptor protein). In addition, the specimen slices can include a counterstain (e.g., hematoxylin). Color deconvolution can be applied to generate a composite image of each biomarker. That is, a first color vector can be applied to the dual IHC image to generate a first composite image depicting a first biomarker based on the color of the first stain, and a second color vector can be applied to the dual IHC image to generate a second composite image depicting a second biomarker based on the color of the second stain. Figures 2 and 3 show examples of dual IHC images 202 / 302 of slices of a specimen stained for estrogen receptor protein and progesterone receptor protein. In Figure 3, dual IHC image 302 is part of an entire slide image 301. Color deconvolution is performed on dual IHC images 202 / 302 to generate composite images 204 / 206 / 304 / 306. Composite images 304 / 404 depict estrogen receptor protein, while composite images 206 / 306 depict progesterone receptor protein, respectively.

[0052] Color deconvolution can be additionally applied to each composite image to generate an image representing intensity (e.g., a grayscale image with pixel values ​​from 0 to 255 representing intensity). Color deconvolution can include determining a stain reference vector from the composite or non-counter-stained image, performing a matrix inversion using the reference vector to determine the contribution of each stain to the optical density or intensity of that pixel, and generating an intensity composite singlet image by recombining the unmixed images.

[0053] FIG. 4 shows an example of an intensity composite image generated by applying color deconvolution to a dual IHC image. For example, for dual IHC image 402A, image 412 represents hematoxylin intensity, image 414A represents Dabcyl estrogen receptor intensity, and image 416A represents Tamra progesterone receptor intensity. Additionally, for dual IHC images 402B-C, images 414B-C represent Dabcyl estrogen receptor intensity, and images 416B-C represent Tamra progesterone receptor intensity, respectively. Dual IHC image 402B can correspond to moderate to strong staining, while dual IHC image 402C can correspond to weak staining.

[0054] Returning to FIG. 1 , the training controller 145 can feed the original or preprocessed images (e.g., the dual IHC image and / or each composite image) to another trained machine learning model having an architecture (e.g., U-Net) configured with parameters used and learned during previous training. The other trained machine learning model can generate an output that identifies a first cell representation predicted to depict positive staining for a first biomarker and a second cell representation predicted to depict positive staining for a second biomarker.

[0055] Referring to FIG. 5, the output of another trained machine learning model is shown. Image 522 shows a dual IHC image with predicted biomarker depictions in various colors. For example, cells depicted in red correspond to positive staining for both biomarkers, cells depicted in green correspond to positive staining for a first biomarker (e.g., estrogen receptor protein), cells depicted in blue correspond to positive staining for a second biomarker (e.g., progesterone receptor protein), cells depicted in yellow correspond to negative staining for both biomarkers, and cells depicted in black correspond to other detected cells (e.g., stoma cells). The composite image may be further input into the trained machine learning model to generate images 524 / 526, or predicted biomarker depictions may be extracted from image 522 and overlaid on the composite image to generate images 524 / 526. In image 524, which corresponds to a composite image depicting estrogen receptor protein, cells depicted in red correspond to positive staining for estrogen receptor protein, cells depicted in yellow correspond to negative staining for estrogen receptor protein, and cells depicted in black correspond to other detected cells. In image 526, which corresponds to a composite image depicting progesterone receptor protein, cells depicted in red correspond to positive staining for progesterone receptor protein, cells depicted in yellow correspond to negative staining for progesterone receptor protein, and cells depicted in black correspond to other detected cells.

[0056] Returning to FIG. 1 , once the other trained machine learning models output predicted biomarker depictions, the training controller 145 can generate inputs for the expression level prediction model 140. Expression level prediction can be performed only on cells predicted to depict positive staining for at least one of the biomarkers. Thus, based on the output of the other trained machine learning models, portions of the dual IHC image and / or composite image depicting positive staining for one or more biomarkers can be extracted. For example, in an intensity composite image of a first biomarker, portions predicted to depict positive staining for the first biomarker can be extracted. Additionally, in an intensity composite image of a second biomarker, portions predicted to depict positive staining for the second biomarker can be extracted. Extracting portions can involve defining a patch (e.g., a 5×5 patch) around each portion predicted to contain positively stained cells.

[0057] 6, an image 632 of cells predicted to depict positive staining for a biomarker is shown overlaid on an intensity composite image of the biomarker. Patch 634 is extracted from image 632. Patch 634 is a 5x5 patch of a portion of image 632 predicted to show cells with positive staining for the biomarker. Multiple patches can be extracted from image 632, and each patch can be a 5x5 patch surrounding a cell predicted to depict positive staining for the biomarker.

[0058] Similarly, FIG. 7 shows image 722 with predicted biomarker depictions and extracted patches 730. Intensity composite images 732A-C show cell depictions predicted to contain positive staining for the biomarkers in image 722, and patches 734A-C show cell depictions predicted to contain positive staining for the biomarkers in patch 730. Image 732A and patch 734A show depicted cells predicted to contain positive staining for both estrogen receptor protein and progesterone receptor protein. Image 732B and patch 734B show depicted cells predicted to contain positive staining in the Dabcyl channel; only cell patches around cells predicted to contain positive staining for estrogen receptor protein are calculated; positive staining progesterone receptor cells are not considered for calculating expression levels. Image 732C and patch 734C show delineated cells predicted to contain positive staining in the Tamra channel, where only cell patches around cells predicted to contain positive staining for progesterone receptor protein are calculated, and positive staining estrogen receptor cells are not considered to calculate expression levels.

[0059] Returning to FIG. 1 , in an example, the training controller 145 can perform feature extraction on each patch to generate a set of features representing pixel intensities of cell depictions in the first composite image and the second composite image. In a first feature extraction technique, the training controller 145 can determine, for each cell in the intensity composite image that is predicted to depict positive staining for the first biomarker, a metric associated with the intensity value of the patch containing the cell. For example, the metric can be the average intensity value of the pixels in the patch. The training controller 145 can then aggregate the metrics for each patch in the intensity composite image. Aggregating the metrics can include ranking the average value of each patch from lowest intensity (e.g., close to 0) to highest intensity (e.g., close to 255). The metrics can additionally be normalized so that each value is between 0 and 1. From the aggregated metrics, the training controller 145 can determine intensity values ​​of cells in the intensity composite image that are predicted to depict positive staining for the first biomarker. Each intensity value can correspond to an intensity percentile derived from the normalized patch intensities. The training controller 145 can perform a similar process for each cell in the intensity composite image that is predicted to depict positive staining for the second biomarker.

[0060] 8A-8B, an exemplary table of intensity percentile intensity values ​​for each of weakly stained image 802A and medium-highly stained image 802B for estrogen receptor protein is shown. Images 804A-B are intensity composite images corresponding to weakly stained image 802A and medium-highly stained image 802B, respectively. The intensity values ​​are aggregated normalized intensity values ​​for each of images 804A-B. For each of images 804A-B, aggregated normalized intensity values ​​are shown for the 10% intensity percentile, 25% intensity percentile, 50% intensity percentile, 90% intensity percentile, 95% intensity percentile, 97.5% intensity percentile, 99% intensity percentile, 99.25% intensity percentile, and 99.5% intensity percentile. For image 804A, the aggregated normalized intensity value for the 10% intensity percentile is 0.326535 and increases with each intensity percentile to 0.763664 for the 99.5% intensity percentile. In contrast, for image 804B, the aggregated normalized intensity value for the 10% intensity percentile is 0.386133 and increases with each intensity percentile to 0.868026 for the 99.5% intensity percentile. At each intensity percentile, the intensity value is greater for medium-high staining image 802B than for weakly stained image 802A.

[0061] Returning to FIG. 1 , an alternative feature extraction technique may include training controller 145 determining, for each patch of the intensity composite image predicted to depict positive staining for the first biomarker, intensity values ​​corresponding to the patch's intensity percentiles (e.g., 50%, 60%, 70%, 80%, 90%, and 95%). Thus, for a given patch, training controller 145 can determine a distribution of intensity values ​​in the patch. Based on the distribution, training controller 145 can determine an intensity value associated with each intensity percentile. Training controller 145 can then aggregate the intensity values ​​of the intensity percentiles for each patch of the intensity composite image. That is, intensity values ​​associated with the 50% percentile of each patch can be aggregated, and intensity values ​​associated with the 60% percentile can be aggregated. Training controller 145 can then calculate a histogram of each intensity percentile and normalize the bins of each histogram.

[0062] Turning to Figure 9, example histograms are shown corresponding to a weakly stained image 904A and a moderately stained image 904B of estrogen receptor protein. The weakly stained image 904A is generated from the double IHC image 902A, and the moderately stained image 904B is generated from the double image 902B. For the weakly stained image, the majority of the intensity values ​​are between 0.3 and 0.7, and for the moderately stained image, the majority of the intensity values ​​are between 0.6 and 0.9.

[0063] 10 shows exemplary histograms corresponding to a weakly stained image 1004A and a moderately stained image 1004B of progesterone receptor protein. The weakly stained image 1004A is generated from the double IHC image 1002A, and the moderately stained image 1004B is generated from the double IHC image 1002B. For the weakly stained image 1004A, the majority of the intensity values ​​are between 0.3 and 0.7, and for the moderately stained image 1004B, the majority of the intensity values ​​are between 0.6 and 0.9.

[0064] 11A-11B show exemplary histograms corresponding to a weakly stained image 1104A and a medium-highly stained image 1104B of estrogen receptor protein. The histograms represent aggregated intensity values ​​for multiple intensity percentiles for the weakly stained image 1104A and the medium-highly stained image 1104B. The intensity percentiles include the 50% intensity percentile, the 60% intensity percentile, the 70% intensity percentile, the 80% intensity percentile, the 90% intensity percentile, and the 95% intensity percentile. For each intensity percentile, the aggregated intensity value is greater for the medium-highly stained image 1104B compared to the weakly stained image 1104A.

[0065] 12A-12B show exemplary histograms corresponding to a weakly stained image 1204A and a medium-highly stained image 1204B of progesterone receptor protein. The histograms represent aggregated intensity values ​​for multiple intensity percentiles for the weakly stained image 1204A and the medium-highly stained image 1204B. The intensity percentiles include the 50% intensity percentile, the 60% intensity percentile, the 70% intensity percentile, the 80% intensity percentile, the 90% intensity percentile, and the 95% intensity percentile. As in FIGS. 11A-11B, for each intensity percentile, the aggregated intensity value is greater for the medium-highly stained image 1204B compared to the weakly stained image 1204A.

[0066] Returning to FIG. 1 , the computing system 100 may include a label mapper 160 that maps an image 130 from the imaging system 125, including depictions of biomarkers associated with a disease, to a label indicating the expression level of the biomarker. The label may be determined based on a pathologist's determination of one or more expression levels of the image 130. For example, one or more pathologists may provide an H-score determination corresponding to the expression level of the biomarker in the image, and the H-score may be used as the label. The H-score may be obtained using the formula: 3 × percentage of strongly stained nuclei + 2 × percentage of moderately stained nuclei + percentage of weakly stained nuclei. If multiple pathologists provide H-scores, the label may be the mean or median H-score among the pathologists. The label may also include intensity values ​​determined from feature extraction. For example, for a first feature extraction technique, the label may include intensity values ​​and corresponding intensity percentiles. For a second feature extraction method, the label may include a distribution of intensity values ​​for each intensity percentile. The mapping data may be stored in a mapping data store (not shown). The mapping data may identify the expression level mapped to each image.

[0067] In some cases, labels associated with training dataset 150 may be received or derived from data received from remote system 155. The received data may include (for example) one or more medical records corresponding to the particular subject to which one or more of images 130 correspond. In some cases, the images or scans input to one or more classifier subsystems are received from remote system 155. For example, remote system 155 may receive images 130 from image generation system 105 and then transmit images 130 or scans (e.g., along with an identifier and one or more labels for the subject) to analysis system 135.

[0068] The training controller 145 can use the mapping of the training dataset 150 to train the expression level prediction model 140. More specifically, the training controller 145 can access the model's architecture, define the model's (fixed) hyperparameters (these are parameters that affect the learning process, such as, for example, the learning rate, the size / complexity of the model, etc.), and train the model so that a set of parameters is learned. More specifically, the set of parameters can be learned by identifying parameter values ​​associated with low or minimal loss, cost, or error generated by comparing predicted outputs (obtained using given parameter values) with actual outputs. In some cases, the machine-learning model can be configured to iteratively fit new models to improve the estimated accuracy of outputs (e.g., including metrics or identifiers corresponding to biomarker expression level predictions).

[0069] The machine learning (ML) execution handler 165 can process independent data using the architecture and learned parameters to generate results. For example, the ML execution handler 165 can access dual IHC images that are not represented in the training dataset 150. In some embodiments, the generated dual IHC images are stored in a memory device. The images can be generated using the imaging system 125. In some embodiments, the images are generated or acquired from a microscope or other instrument capable of capturing image data of a microscope slide holding a specimen, as described herein. In some embodiments, the biomedical images are generated or acquired using a 2D scanner, such as one capable of scanning image tiles. Alternatively, the images can be previously generated (e.g., scanned) and stored in a memory device (or, in that case, retrieved from a server via a communication network).

[0070] In some cases, the dual IHC images may be preprocessed according to learned or identified preprocessing techniques. For example, the ML execution handler 165 may apply color deconvolution to the dual IHC images to generate a composite image depicting each of the biomarkers. Furthermore, the ML execution handler 165 may apply additional color deconvolution to each composite image to generate an intensity composite image. The original and / or preprocessed images (e.g., the dual IHC images and / or each composite image) may be fed to a trained machine learning model having an architecture (e.g., U-Net) configured with learned parameters used during training. The trained machine learning model may generate an output that distinguishes between a first cell representation predicted to depict a first biomarker and a second cell representation predicted to depict a second biomarker.

[0071] Once the trained machine learning model outputs predicted biomarker depictions, the ML execution handler 165 can predict the expression levels of the biomarkers using the architecture and learned parameters of the expression level prediction model 140. Expression level prediction can be performed only on cells predicted to depict positive staining for at least one of the biomarkers. Thus, based on the output of the trained machine learning model, portions of the dual IHC image and / or composite image depicting positive staining for one or more biomarkers can be extracted. For example, in an intensity composite image of a first biomarker, portions predicted to depict positive staining for the first biomarker can be extracted. Additionally, in an intensity composite image of a first biomarker, portions predicted to depict positive staining for the first biomarker can be extracted. Extracting portions can involve defining patches (e.g., 5x5 patches) around each portion predicted to contain positively stained cells. The ML execution handler 165 can then perform feature extraction techniques on the intensity composite image to determine intensity values ​​associated with the intensity percentiles of each patch and the entire image.

[0072] The original and / or preprocessed images (e.g., dual IHC images, each composite image, and / or each intensity composite image) and intensity values ​​may be fed to an expression level prediction model 140 having an architecture (e.g., a linear regression model) configured with parameters used and learned during training. The expression level prediction model 140 may generate an output identifying predicted expression levels of the first biomarker and the second biomarker.

[0073] In some cases, the image characteristic evaluator 170 identifies a predicted characteristic of disease for the image based on the execution of image processing. The execution of the expression level prediction model 140 itself may generate a result including a characteristic, or the execution may include a result that the image characteristic evaluator 170 can use to determine a predicted characteristic of the specimen. For example, the image characteristic evaluator 170 may perform subsequent processing that may include characterizing the presence, quantity, and / or size of a set of tumor cells predicted to be present in the image. The subsequent processing may additionally or alternatively include characterizing the diagnosis of a disease predicted to be present in the image, classifying a disease predicted to be present in the image, and / or predicting a prognosis of a disease predicted to be present in the image. The image characteristic evaluator 170 may apply rules and / or transformations to map the predicted expression levels and associated probabilities and / or confidences to a characteristic. Illustratively, if the result includes a probability of the predicted expression level being above a threshold that is greater than 50%, a first characteristic may be assigned; otherwise, a second characteristic may be assigned.

[0074] Communications interface 175 can collect the results and communicate the results (or a processed version thereof) to a user device (e.g., associated with a laboratory technician or care provider) or other system. For example, the results may be communicated to remote system 155. In some cases, communications interface 175 may generate an output that identifies the presence, amount, and / or size of a set of tumor cells, a disease diagnosis, a disease classification, and / or a disease prognosis. The output may then be presented and / or transmitted, which may facilitate display of the output data, for example, on a display of a computing device. The results may be used to determine a diagnosis, a treatment plan, or to evaluate ongoing treatment of the tumor cells.

[0075] III. Exemplary Use Cases 13 shows an exemplary process for predicting expression levels in digital pathology images. The steps of the process may be performed by one or more systems. Other examples may include more steps, fewer steps, different steps, or steps in a different order.

[0076] In block 1305, a dual IHC image of a slice of the specimen is accessed. The dual IHC image can include cellular depictions associated with one or more of a first biomarker and a second biomarker corresponding to a disease. For example, to identify breast cancer, the first biomarker can be estrogen receptor protein and the second biomarker can be progesterone receptor protein. The slice of the specimen can include a first stain for the first biomarker and a second stain for the second biomarker. By way of example, the first stain can be Dabcyl and the second stain can be Tamra.

[0077] At block 1310, a first composite image and a second composite image are generated. Color deconvolution can be applied to the dual IHC images to generate the first composite image and the second composite image. The first composite image can depict a first biomarker, and the second composite image can depict a second biomarker. Additional preprocessing can also be applied to the composite images. For example, additional color deconvolution can be applied to the first composite image and the second composite image to generate an intensity composite image with grayscale pixels representing the intensity of cellular depictions in the composite image. The composite images can also be input to a trained machine learning model that distinguishes between cellular depictions in the first composite image that are predicted to depict a first biomarker and cellular depictions in the second composite image that are predicted to depict a second biomarker.

[0078] In block 1315, a set of features representing pixel intensities of cell representations is determined. Patches may be generated, each containing at least one cell representation predicted to depict either a first biomarker or a second biomarker. For each cell predicted to depict positive staining for the first biomarker, a metric associated with the intensity value of the patch containing the cell may be determined. For example, the metric may be the average intensity value of the pixels in the patch. The metrics for each patch in the intensity composite image may then be aggregated and normalized. From the aggregated metric, an intensity value may be determined for cells in the intensity composite image predicted to depict positive staining for the first biomarker. Each intensity value may correspond to an intensity percentile derived from the normalized patch intensity. A similar process may be performed for each cell in the intensity composite image predicted to depict positive staining for the second biomarker. An alternative feature extraction technique may include determining, for each patch of the intensity composite image predicted to depict positive staining for the first biomarker, intensity values ​​corresponding to the patch's intensity percentiles (e.g., 50%, 60%, 70%, 80%, 90%, and 95%). The intensity values ​​associated with each intensity percentile may be determined, and the intensity values ​​for the intensity percentiles for each patch of the intensity composite image may be aggregated. A set of metrics associated with the distribution of the aggregated intensity values ​​for the intensity percentiles may be determined. For example, the set of metrics may be determined from a histogram generated for each intensity percentile.

[0079] At block 1320, the set of features is processed using the trained machine learning model. For a first feature extraction technique, the set of features may be intensity values ​​corresponding to different intensity percentiles. For a second feature extraction technique, the set of features may be a set of metrics associated with the distribution of aggregated intensity values ​​relative to the intensity percentiles.

[0080] At block 1325, results corresponding to a predicted characterization of the specimen with respect to the disease are output. For example, the results may be transmitted to and / or displayed on another device (e.g., associated with a care provider). The results may correspond to a predicted characterization of the specimen. The results may characterize the presence, amount, and / or size of a set of tumor cells in the image, a disease diagnosis, a disease classification, and / or a disease prognosis.

[0081] IV. Illustrative Results Figures 14A-14C show exemplary expression levels of Dabcyl estrogen receptor protein determined by three pathologists for 50 fields (e.g., dual IHC images). In scoring Dabcyl estrogen receptor protein, examples of moderate to high staining had high consistency across pathologists, while examples of low staining had greater variance among the three pathologists. The expression levels determined by each of the three pathologists were compared to the median estrogen receptor protein expression levels across pathologists for each field. As shown, there was high consistency across the three pathologists, with correlation coefficients between 0.93 and 0.98, respectively.

[0082] Figures 15A-15C show exemplary expression levels of Tamra progesterone receptor protein as determined by three pathologists for 50 fields (e.g., dual IHC images). In scoring Tamra progesterone receptor protein, examples of moderate to high staining had high consistency across pathologists, while examples of low staining had more variance among the three pathologists. The expression levels determined by each of the three pathologists were compared to the median progesterone receptor protein expression levels across pathologists for each field. As shown, there was high consistency across the three pathologists, with correlation coefficients between 0.88 and 0.96, respectively.

[0083] Figures 16A-16B show exemplary performance of predicting expression levels using a machine learning model. Intensity values ​​were extracted using the second feature extraction technique described herein. The trained machine learning model (e.g., expression level prediction model 140 in Figure 1) achieved high consistency in predicting expression levels compared to the median estrogen receptor expression levels determined by pathologists and the median progesterone receptor expression levels determined by pathologists. Compared to the median consensus expression level, the trained machine learning model achieved higher consistency in predicting expression levels than the scoring performed by pathologists. The correlation was 0.9788 for the estrogen receptor protein expression level prediction by the trained machine learning model compared to the median consensus expression level, and 0.9292 for the progesterone receptor protein expression level prediction by the trained machine learning model compared to the median consensus expression level. The predicted R-squared was 0.958 and 0.8635, respectively. The table further shows the correlation between the three pathologists and the trained machine learning model. In scoring Dabcyl estrogen receptor protein expression levels, the trained machine learning model achieved a higher correlation with the median consensus than either of the pathologists. Additionally, in scoring Tamra progesterone receptor protein expression levels, the trained machine learning model outperformed two of the pathologists in terms of correlation to the median consensus. As a result, the trained machine learning model can generate more consistently accurate expression level predictions than pathologists, which can facilitate more accurate characterization of disease.

[0084] FIG. 17 shows exemplary expression level scores generated by a pathologist and a trained machine learning model. For image 1702, the pathologist determined an estrogen receptor protein expression level of 2.15, and the trained machine learning model determined an estrogen receptor protein expression level of 2.10. For image 1704, the pathologist determined an estrogen receptor protein expression level of 1.15, and the trained machine learning model determined an estrogen receptor protein expression level of 1.16. For image 1706, the pathologist determined a progesterone receptor protein expression level of 2.40, and the trained machine learning model determined an progesterone receptor protein expression level of 2.35. For image 1708, the pathologist determined a progesterone receptor protein expression level of 1.50, and the trained machine learning model determined an progesterone receptor protein expression level of 1.53. In each case, the predictions by the trained machine learning model were within 0.05 of the scores determined by the pathologists, further demonstrating the accuracy of the trained machine learning model in predicting biomarker expression levels in digital pathology images.

[0085] V. Further Considerations Some embodiments of the present disclosure include a system including one or more data processors. In some embodiments, the system includes a non-transitory computer-readable storage medium including instructions that, when executed on the one or more data processors, cause the one or more data processors to perform some or all of one or more methods and / or some or all of one or more processes disclosed herein. Some embodiments of the present disclosure include a computer program product tangibly embodied in a non-transitory machine-readable storage medium, the computer program product including instructions configured to cause one or more data processors to perform some or all of one or more methods and / or some or all of one or more processes disclosed herein.

[0086] The terms and expressions which have been employed are used as terms of description rather than of limitation, and there is no intention in the use of such terms and expressions to exclude equivalents of the features shown and described or portions thereof, but it is recognized that various modifications are possible within the scope of the invention as claimed. Thus, although the claimed invention has been specifically disclosed by embodiments and optional features, it is to be understood that modifications and variations of the concepts disclosed herein may be employed by those skilled in the art, and that such modifications and variations are deemed to be within the scope of the invention as defined by the appended claims.

[0087] The description presents only preferred exemplary embodiments and is not intended to limit the scope, applicability, or configuration of the present disclosure. Rather, the description of preferred exemplary embodiments provides those skilled in the art with an enabling description for implementing various embodiments. It will be understood that various changes can be made in the function and arrangement of elements without departing from the spirit and scope as set forth in the appended claims.

[0088] In the following description, specific details are given to provide a comprehensive understanding of the embodiments. However, it will be understood that the embodiments may be practiced without these specific details. For example, circuits, systems, networks, processes, and other components may be shown as components in block diagram form in order to avoid obscuring the embodiments in unnecessary detail. In other instances, well-known circuits, processes, algorithms, structures, and techniques may be shown without unnecessary detail in order to avoid obscuring the embodiments.

Claims

1. 1. A computer-implemented method comprising: accessing a dual immunohistochemistry (IHC) image of a slice of the specimen, the dual IHC image including cellular delineations associated with one or more of a first biomarker and a second biomarker corresponding to the disease; generating a first composite image depicting the first biomarker and a second composite image depicting the second biomarker from the dual IHC image; determining, for each of the first and second composite images, a set of features representative of pixel intensities of the cell depictions in the first and second composite images; processing the set of features using a trained machine learning model, wherein an output of the processing corresponds to a predicted expression level of the first biomarker and the second biomarker; and outputting a result corresponding to a predicted characterization of the specimen with respect to the disease based on the output of the process.

2. before determining the set of features, 2. The computer-implemented method of claim 1, further comprising preprocessing the first and second composite images by applying color deconvolution to the first and second composite images.

3. before determining the set of features, 2. The computer-implemented method of claim 1, further comprising processing the first composite image and the second composite image using separate trained machine learning models, wherein separate outputs of the processing identify first cell representations in the first composite image that are predicted to depict the first biomarker and second cell representations in the second composite image that are predicted to depict the second biomarker.

4. Determining the set of features of the first composite image comprises: determining, for each cell of the first cell representation, a first metric associated with an intensity value of a patch of the cell that includes the cell; aggregating the first metric for each patch for the first cell representation; and 4. The computer-implemented method of claim 3, comprising determining a plurality of intensity values ​​for the first cell representation based on the aggregation, wherein each intensity value of the plurality of intensity values ​​corresponds to an intensity percentile, and wherein the plurality of intensity values ​​corresponds to the set of features.

5. Determining the set of features of the first composite image comprises: determining, for each cell in the first cell representation, a first plurality of intensity values ​​corresponding to intensity percentiles of a patch containing the cell; aggregating the first plurality of intensity values ​​for each patch for the first cell representation to generate a second plurality of intensity values; and 4. The computer-implemented method of claim 3, comprising determining a set of metrics associated with a distribution of the second plurality of intensity values, the set of metrics corresponding to the set of features.

6. 10. The computer-implemented method of claim 1, wherein the first biomarker comprises an estrogen receptor protein and the second biomarker comprises a progesterone receptor protein.

7. The method of claim 1 , wherein the trained machine learning model comprises a linear regression model.

8. 10. The computer-implemented method of claim 1, wherein a sample slice of the specimen comprises a first stain for the first biomarker and a second stain for the second biomarker.

9. 9. The computer-implemented method of claim 8, wherein the first stain comprises tetramethylrhodamine and the second stain comprises 4-dimethylaminoazobenzene-4'-sulfonyl.

10. 2. The computer-implemented method of claim 1, further comprising performing subsequent processing to generate the result of the predicted characterization of the specimen, wherein performing the subsequent image processing comprises detecting a delineation of a set of tumor cells, and wherein the result characterizes the presence, amount, and / or size of the set of tumor cells.

11. 1. A system comprising: one or more data processors; a non-transitory computer-readable storage medium comprising instructions that, when executed on the one or more data processors, cause the one or more data processors to perform operations, the operations including: accessing a dual immunohistochemistry (IHC) image of a slice of the specimen, the dual IHC image including cellular delineations associated with one or more of a first biomarker and a second biomarker corresponding to the disease; generating a first composite image depicting the first biomarker and a second composite image depicting the second biomarker from the dual IHC image; determining, for each of the first and second composite images, a set of features representative of pixel intensities of the cell depictions in the first and second composite images; processing the set of features using a trained machine learning model, wherein an output of the processing corresponds to a predicted expression level of the first biomarker and the second biomarker; and outputting a result corresponding to a predicted characterization of the specimen with respect to the disease based on the output of the processing.

12. The non-transitory computer-readable storage medium further includes instructions that, when executed by the one or more data processors, cause the one or more data processors to, before determining the set of characteristics:

12. The system of claim 11, further comprising: preprocessing the first and second composite images by applying color deconvolution to the first and second composite images.

13. The non-transitory computer-readable storage medium further includes instructions that, when executed by the one or more data processors, cause the one or more data processors to, before determining the set of characteristics:

12. The system of claim 11, further comprising: processing the first composite image and the second composite image using separate trained machine learning models, wherein separate outputs of the processing identify first cellular representations in the first composite image that are predicted to depict the first biomarker and second cellular representations in the second composite image that are predicted to depict the second biomarker.

14. Determining the set of features of the first composite image comprises: determining, for each cell of the first cell representation, a first metric associated with an intensity value of a patch of the cell that includes the cell; aggregating the first metric for each patch for the first cell representation; and 14. The system of claim 13, further comprising determining a plurality of intensity values ​​for the first cell representation based on the aggregation, wherein each intensity value of the plurality of intensity values ​​corresponds to an intensity percentile, and wherein the plurality of intensity values ​​corresponds to the set of features.

15. Determining the set of features of the first composite image comprises: determining, for each cell in the first cell representation, a first plurality of intensity values ​​corresponding to intensity percentiles of a patch containing the cell; aggregating the first plurality of intensity values ​​for each patch for the first cell representation to generate a second plurality of intensity values; and 14. The system of claim 13, further comprising: determining a set of metrics associated with a distribution of the second plurality of intensity values, the set of metrics corresponding to the set of features.

16. 14. The system of claim 13, wherein the first biomarker comprises an estrogen receptor protein and the second biomarker comprises a progesterone receptor protein.

17. The system of claim 11 , wherein the trained machine learning model comprises a linear regression model.

18. 12. The system of claim 11, wherein a sample slice of the specimen comprises a first stain for the first biomarker and a second stain for the second biomarker.

19. 19. The system of claim 18, wherein the first stain comprises tetramethylrhodamine and the second stain comprises 4-dimethylaminoazobenzene-4'-sulfonyl.

20. The non-transitory computer-readable storage medium further includes instructions that, when executed by the one or more data processors, cause the one or more data processors to:

12. The system of claim 11, further comprising: performing a subsequent process to generate the result of the predicted characterization of the specimen, wherein performing the subsequent process comprises detecting a delineation of a set of tumor cells, and wherein the result causes an operation to be performed that characterizes the presence, amount, and / or size of the set of tumor cells.

21. A computer program product tangibly embodied in a non-transitory machine-readable storage medium comprising instructions configured to cause one or more data processors to perform operations, said operations comprising: accessing a dual immunohistochemistry (IHC) image of a slice of the specimen, the dual IHC image including cellular delineations associated with one or more of a first biomarker and a second biomarker corresponding to the disease; generating a first composite image depicting the first biomarker and a second composite image depicting the second biomarker from the dual IHC image; determining, for each of the first and second composite images, a set of features representative of pixel intensities of the cell depictions in the first and second composite images; processing the set of features using a trained machine learning model, wherein an output of the processing corresponds to a predicted expression level of the first biomarker and the second biomarker; and outputting a result corresponding to a predicted characterization of the specimen with respect to the disease based on the output of the process.