Method for determining biomarkers from pathological tissue slide images
A deep learning framework efficiently analyzes histopathology images to identify biomarkers like TILs and PD-L1, addressing inefficiencies in current methods by using multi-scale and single-scale configurations, enhancing treatment recommendations and disease prediction.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2025-05-14
- Publication Date
- 2026-03-10
AI Technical Summary
Current methods for analyzing tumor samples, such as immunohistochemistry (IHC) staining, are limited by insufficient tissue samples and resource constraints, and existing deep learning approaches like CNNs and FCNs are inefficient for classifying large histopathology images due to redundant computation and annotation requirements.
A deep learning framework is developed to analyze histopathology images directly, using multi-scale and single-scale configurations with classifiers trained on unlabeled or labeled images, incorporating tile-level and pixel-level cell segmentation to predict biomarkers like TILs and PD-L1, reducing computation time and resource requirements.
The framework efficiently identifies biomarkers across various cancers, enabling optimized treatment recommendations and improving disease progression prediction by automating the analysis of large histopathology images with reduced computational time and resource needs.
Smart Images

Figure 0007827919000001 
Figure 0007827919000002 
Figure 0007827919000003
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application is a continuation-in-part of U.S. patent application Ser. No. 16 / 732,242, filed December 31, 2019, which claims priority to U.S. provisional patent application Ser. No. 62 / 787,047, filed December 31, 2018, and is a continuation-in-part of U.S. patent application Ser. No. 16 / 41, filed May 14, 2019, which claims priority to U.S. provisional patent application Ser. No. 62 / 671,300, filed May 14, 2018. This application is a continuation-in-part of U.S. Provisional Patent Application No. 2,362, which claims priority to U.S. Provisional Patent Application No. 62 / 824,039, filed March 26, 2019, U.S. Provisional Patent Application No. 62 / 889,521, filed August 20, 2019, and U.S. Provisional Patent Application No. 62 / 983,524, filed February 28, 2020, the entire disclosures of each of which are expressly incorporated herein by reference.
[0002] The present disclosure relates to detecting, quantifying, and / or characterizing biomarkers associated with cancer, and more particularly to examining digital images to detect, quantify, and / or characterize such biomarkers from the analysis of one or more histopathology slide images. [Background technology]
[0003] The description of the background art provided herein is for the purpose of generally presenting the context of the present disclosure. To the extent that it is described in this background art section, the work of the presently named inventors, and aspects of the description that may not otherwise be considered prior art at the time of filing, are not admitted, expressly or impliedly, to be prior art to the present disclosure.
[0004] Tumor samples are commonly extracted and examined from patients to guide medical professionals in diagnosing, prognosing, and evaluating treatment for a patient's cancer. Visual inspection can reveal the growth pattern of cancer cells within the tumor in relation to nearby healthy cells, as well as the presence of immune cells within the tumor. Traditionally, a pathologist, a member of a pathology team, another trained medical professional, or another human analyst visually analyzes thin slices of tumor tissue mounted on glass microscope slides to identify each region of tissue that corresponds to one of the many tissue types present in the tumor sample. This information helps the pathologist determine the characteristics of the patient's cancer tumor and may aid in treatment decisions. Pathologists often assign one or more numerical scores to the slides based on their visual approximation.
[0005] To make these visual approximations, medical professionals attempt to identify many features of the tumor, including, for example, the tumor's grade, the tumor's purity, the degree of invasiveness of the tumor, the degree of immune infiltration of the tumor, the cancer stage, and the tumor's anatomical site of origin, which may be important in diagnosing and treating metastatic tumors. These details about the cancer can help physicians monitor the progression of the cancer within a patient and predict which anti-cancer treatments are likely to be successful in eliminating cancer cells from the patient's body.
[0006] Another characteristic of tumors is the presence of specific biomarkers or other cell types, including immune cells, within or near tumors. For example, high levels of tumor-infiltrating lymphocytes (TILs) have been recognized as a biomarker of antitumor immune responses across a wide range of tumors. TILs are mononuclear immune cells that infiltrate tumor tissue or stroma and have been reported in multiple tumor types, including breast cancer. TIL populations are composed of various cell types (i.e., T cells, B cells, natural killer (NK) cells, etc.). While naturally occurring TIL populations in cancer patients are largely ineffective at destroying tumors, the presence of TILs has been associated with improved prognosis in many types of cancer, such as epithelial ovarian cancer, colon cancer, esophageal cancer, melanoma, endometrial cancer, and breast cancer (see, e.g., Melichar et al., Anticancer Res. 2014;34(3):1115-25; Naito et al., Cancer Res. 1998;58(16):3491-4).
[0007] Another characteristic of tumors is the presence of specific molecules known as biomarkers, including the molecule known as programmed death-ligand 1 (PD-L1). PD-L1 is relevant to the diagnosis and evaluation of non-small cell lung cancer (NSCLC). NSCLC is the most common type of lung cancer, affecting more than 1.5 million people worldwide. NSCLC has a poor response to standard treatment, chemoradiotherapy, and a high incidence of recurrence, resulting in a low 5-year survival rate. Advances in immunology have revealed that NSCLC frequently expresses elevated levels of programmed death-1 (PD-L1), which binds to PD-1 on the surface of T cells. Binding of PD-1 to PD-L1 deactivates T cell antitumor responses, allowing NSCLC to evade immune system targeting. The discovery of the interplay between tumor progression and the immune response has led to the development and regulatory approval of PD-1 / PD-L1 checkpoint blockade immunotherapies, such as nivolumab and pembrolizumab. Anti-PD-1 and anti-PD-L1 antibodies restore antitumor immune responses by disrupting the interaction between PD-1 and PD-L1. Notably, patients with PD-L1-positive NSCLC treated with these checkpoint inhibitors have achieved durable tumor regression and improved survival.
[0008] As the role of immunotherapy in oncology expands, accurately assessing tumor PD-L1 status may help identify patients who may benefit from PD-1 / PD-L1 checkpoint blockade immunotherapy. Currently, immunohistochemistry (IHC) staining of tumor tissue obtained from biopsies or surgical specimens is employed to assess PD-L1 status. However, such IHC staining is often limited by insufficient tissue samples or limited resources in some settings.
[0009] Hematoxylin and eosin (H&E) staining is a long-standing method used by pathologists to analyze the morphological characteristics of tissue for the diagnosis of malignant tumors. For example, H&E slides can reveal visual characteristics of tissue structures, such as cell nuclei and cytoplasm, to aid in the identification of cancerous tumors.
[0010] Technological advances have made it possible to digitize histopathological H&E and IHC slides into high-resolution whole slide images (WSIs), providing opportunities for developing computer vision tools for a wide range of clinical applications. High-resolution digital images of microscope slides allow for computer-assisted analysis of the slides, enabling tissue type or pathology classification. Broadly speaking, deep learning applications have shown promise, for example, as a tool in medical diagnostic applications and treatment outcome prediction. Deep learning is a subset of machine learning, and models can be constructed with multiple individual neural node layers. A convolutional neural network (CNN) is a neural network that employs convolutional techniques. For example, a CNN can provide a deep learning process, analyzing digital images by assigning a class label to each input image. However, WSIs contain two or more types of tissue, including boundaries between adjacent tissue classes. To analyze the boundaries between adjacent tissue classes and the presence of immune cells between tumor cells, it is necessary to classify partially distinct regions as different tissue classes. For a traditional CNN to assign multiple tissue classes to a single slide image, the CNN must separately process each section of the image that needs to be assigned a tissue class label. However, because sections of adjacent images overlap, processing each section separately requires a lot of redundant computation and is time-consuming.
[0011] A fully convolutional network (FCN) is another type of deep learning process. An FCN can analyze an image and assign a classification label to each pixel in the image. As a result, compared to a CNN, an FCN is useful for analyzing images that represent objects with two or more classifications. An FCN generates an overlay map showing the location of each classified object in the original image. However, for an FCN deep learning algorithm to be effective, it must be trained on a dataset of images in which each pixel is labeled as a tissue class, which requires too much annotation and processing time to be practical. Digital WSI images can contain 10,000 to over 100,000 pixels on each edge of the image. A complete image requires at least 10,000 2 ~100,000 2 A slide may contain pixels, which would require very long algorithm run times to attempt tissue classification. The large number of pixels makes it impossible to segment digital images of slides using traditional FCNs.
[0012] New technologies that can easily diagnose TILs, PD-L1, and other biomarkers using H&E images are needed to identify and characterize such biomarkers across population groups in an efficient manner, to generate more optimized drug treatment recommendations and protocols, and to improve prediction of disease progression. Summary of the Invention
[0013] This application presents an imaging-based biomarker prediction system formed with a deep learning framework configured and trained to learn directly from histopathology slide images and predict the presence of biomarkers in medical images. In examples, the deep learning framework is configured and trained to analyze histopathology images and identify multiple different biomarkers. In various examples, these deep learning frameworks are configured to include different trained biomarker classifiers, each configured to receive unlabeled histopathology images and provide a different biomarker prediction for those images. These biomarker predictions can then be used to reduce a large set of available immunotherapies to a reduced, smaller subset of targeted immunotherapies that medical professionals can use to treat patients. Thus, in various examples, a deep learning framework is provided that identifies biomarkers indicative of the presence of a tumor, the state / condition of the tumor, or information about the tumor in a tissue sample, from which a set of targeted immunotherapies can be determined.
[0014] In an example, the system includes a deep learning framework that is trained to analyze and predict biomarker status in histopathology images received from a network-accessible image source, such as a medical laboratory or medical imaging machine, and generate, store, and display reports of the predicted biomarker status. These predicted biomarker status reports can be provided and stored and displayed in a network-accessible system, such as a pathology laboratory or primary care physician system, and used to determine a patient's cancer treatment protocol (i.e., immunotherapy or chemotherapy treatment). In some examples, the predicted biomarker status reports can be input into a network-accessible next-generation sequencing system to drive subsequent genome sequencing, or into a computerized cancer treatment decision system to filter treatment lists to biomarker-determined treatments.
[0015] The technology herein can identify biomarkers associated with any of a wide variety of cancers. Exemplary cancers include, but are not limited to, adrenocortical carcinoma, lymphoma, anal cancer, anorectal cancer, basal cell carcinoma, skin cancer (non-melanoma), cholangiocarcinoma, extrahepatic bile duct cancer, intrahepatic bile duct cancer, bladder cancer, urinary bladder cancer, osteosarcoma, brain tumor, brainstem glioma, breast cancer (including triple-negative breast cancer), cervical cancer, colon cancer, colorectal cancer, lymphoma, endometrial cancer, esophageal cancer, gastric (stomach) cancer, head and neck cancer, hepatocellular (liver) cancer, renal cancer, kidney cancer, lung cancer, melanoma, tongue cancer, oral cancer, ovarian cancer, pancreatic cancer, prostate cancer, uterine cancer, testicular cancer, and vaginal cancer.
[0016] In some examples, the imaging-based biomarker prediction system is formed in a deep learning framework, the deep learning framework having a multi-scale configuration designed to perform classification of the histopathology image (labeled or unlabeled) using a classifier trained to classify tiles of the received histopathology image. In some examples, the multi-scale configuration includes a tile-level tissue classifier, i.e., a classifier trained using tile-based deep learning training. In some examples, the multi-scale configuration includes a pixel-level cell classifier and a cell segmentation model. In some examples, classifications from the tile-level tissue classifier and the pixel-level cell classifier are analyzed to predict the status of biomarkers in the histopathology image. Furthermore, in some examples, the multi-scale configuration includes a tile-level biomarker classifier.
[0017] In some examples, the imaging-based biomarker prediction system is formed within a deep learning framework, which has a single-scale configuration designed to perform classification of (labeled or unlabeled) histopathological images using a classifier trained using multi-instance learning (MIL) techniques. In some examples, the single-scale configuration includes a slide-level classifier trained using gene sequencing data, such as RNA sequencing data. That is, the slide-level classifier is trained using the RNA sequencing data to develop an image-based classifier that can predict the status of biomarkers in the histopathological images.
[0018] According to one example, a computer-implemented method for identifying biomarkers in a digital image of a hematoxylin and eosin (H&E) stained slide of a target tissue includes: receiving the digital image into an image-based biomarker prediction system having one or more processors; performing, using the one or more processors, an image tiling process on the digital image by separating the digital image into a plurality of tile images, each of the plurality of tile images comprising a different portion of the digital image; applying, using the one or more processors, the plurality of tile images to a multi-scale deep learning framework including one or more trained deep learning multi-scale classifier models, each trained to classify a different tissue classification for each tile image, and determining, using the multi-scale deep learning framework, a tissue classification for each of the plurality of tile images; identifying, using the one or more processors, cells in the digital image using the trained cell segmentation model; and identifying, from the tissue classification determined for each tile image and from the identified cells in the digital image, a predicted presence of one or more biomarkers associated with the digital image.
[0019] According to another example, a computer-implemented method for identifying biomarkers in digital images of hematoxylin and eosin (H&E) stained slides of a target tissue includes receiving a molecular training dataset of a plurality of training tissue samples, the molecular training dataset including RNA transcriptome counts from sequencing of substantially similar samples associated with each training tissue sample; performing a clustering process on the molecular training dataset to identify one or more molecular data subsets, each corresponding to a distinct biomarker; and, for each of the one or more molecular data subsets, performing an image-based biomarker analysis with one or more processors. receiving a plurality of digital images of H&E-stained training slides of training tissue samples corresponding to respective biomarkers for the prediction system; generating, using one or more processors, a trained image-based biomarker classifier model for each of the one or more molecular data subsets based on the plurality of digital images of the H&E-stained training slides; receiving, using one or more processors, subsequent digital images of H&E-stained slides of subsequent tissue samples; and applying, using one or more processors, the subsequent digital images to the trained image-based biomarker classifier model to identify a predicted presence of one or more biomarkers in the subsequent tissue samples.
[0020] According to another example, a computer-implemented method for identifying biomarkers in a digital image of a hematoxylin and eosin (H&E) stained slide of a target tissue includes: receiving the digital image into an image-based biomarker prediction system having one or more processors; separating, using the one or more processors, the digital image into a plurality of tile images, each of the plurality of tile images comprising a different portion of the digital image; applying, using the one or more processors, the plurality of tile images to a deep learning framework including one or more trained biomarker classification models, each trained to classify a different tissue classification; predicting, using the one or more processors, a biomarker classification for each of the plurality of tile images using the one or more trained biomarker classification models; determining a predicted presence of one or more biomarkers in the target tissue from the predicted biomarker classification for each of the tile images; and generating a report including the digital image and a digital overlay visualizing the predicted presence of the one or more biomarkers.
[0021] In some examples, the deep learning framework includes a multi-scale deep learning framework.
[0022] In some examples, separating the digital image into a plurality of tile images includes performing, using one or more processors, an image tiling process by applying a tiling mask to the digital image to separate the digital image into a plurality of tile images.
[0023] In some examples, the tiling mask includes tiles of the same size and / or tiles having a rectangular shape.
[0024] In some examples, applying the plurality of tile images to a deep learning framework and predicting a biomarker classification for each of the plurality of tile images each includes applying each of the tile images to one or more trained deep learning multi-scale classifier models, each trained to classify a different tissue classification for each tile image, and determining a tissue classification for each of the plurality of tile images using the multi-scale deep learning framework; using one or more processors, identifying cells in the digital image using the trained cell segmentation model; and predicting a biomarker classification for each tile image from the tissue classification determined for each tile image and from the identified cells in the digital image.
[0025] In some examples, the method further includes training one or more trained deep learning multi-scale classifier models in a multi-scale deep learning framework by receiving a plurality of H&E slide training images from a training image dataset, each H&E slide training image having a label corresponding to a biomarker to be trained on; performing a tile-based tissue classification analysis on each of the H&E slide training images; performing a pixel-based cell segmentation analysis on each of the H&E slide training images; and optionally performing a tile-based biomarker classification analysis on each of the H&E slide training images; and generating one or more trained deep learning multi-scale classifier models accordingly.
[0026] In some examples, each H&E slide training image includes multiple tile images, each with a tile-level label.
[0027] In some examples, the method includes, for each H&E slide training image, labeling each of a plurality of tile images of the H&E slide training image with a tile-level label.
[0028] In some examples, the method further includes performing, for each H&E slide training image, a tile selection process that infers a class status for each tile image in the H&E slide training image, and discarding tile images that do not correspond to a class of interest before performing a tile-based tissue classification analysis on each of the H&E slide training images based on the inferred class status, thereby performing the tile-based tissue classification analysis only on selected tile images of the H&E slide training images.
[0029] In some examples, each of the one or more trained deep learning multi-scale classifier models is configured as a tile-resolution fully convolutional network (FCN) classification model.
[0030] In some examples, identifying cells in the digital image tiles using the trained cell segmentation model includes applying, using one or more processors, each of the plurality of tile images to the cell segmentation model and, for each tile, assigning a cell classification to one or more pixels in the tile image.
[0031] In some examples, assigning a cell classification to one or more pixels in the tile image includes using one or more processors to identify the one or more pixels as a cell interior, a cell boundary, or a cell exterior, and classifying the one or more pixels as a cell interior, a cell boundary, or a cell exterior.
[0032] In some examples, the trained cell segmentation model is a pixel-resolution 3D UNet classification model trained to classify cell interiors, cell boundaries, and cell exteriors.
[0033] In some examples, the one or more biomarkers are selected from the group consisting of tumor infiltrating lymphocytes (TIL), nuclear-cytoplasmic ratio (NC), ploidy, signet ring morphology, and programmed death-ligand 1 (PD-L1).
[0034] In some examples, the deep learning framework includes a single-scale deep learning framework.
[0035] In some examples, separating the digital image into a plurality of tile images includes performing an image tiling process using one or more processors by applying the digital image to a trained multi-instance learning controller that separates the digital image into a plurality of tile images.
[0036] In some examples, the method further includes providing each tile image to a tile selection process that infers a class status for each tile image in the H&E slide training image, and selectively discarding tile images based on the inferred class status based on tile selection criteria before applying the remaining plurality of tile images to the deep learning framework.
[0037] In some examples, the method further includes providing each tile image to a tile selection process that infers a class status for each tile image in the H&E slide training image, and randomly discarding tile images based on the inferred class status before applying the remaining tile images to the deep learning framework.
[0038] In some examples, the method further includes receiving a molecular training dataset of a plurality of training tissue samples, the molecular training dataset including RNA transcriptome counts from sequencing of substantially similar samples associated with each training tissue sample; performing a clustering process on the molecular training dataset to identify one or more molecular data subsets, each corresponding to a distinct biomarker; receiving, for each of the one or more molecular data subsets, a plurality of digital images of H&E-stained training slides of the training tissue samples corresponding to the respective biomarkers for an image-based biomarker prediction system having one or more processors; and generating, using the one or more processors, one of the trained biomarker classification models for each of the one or more molecular data subsets based on the plurality of digital images of the H&E-stained training slides.
[0039] In some examples, generating one of the trained biomarker classification models for each of the one or more molecular data subsets includes performing a multi-instance learning process on a plurality of digital images of H&E stained training slides.
[0040] In some examples, each of the plurality of digital images of H&E stained training slides of training tissue samples has a slide-level label.
[0041] In some examples, each of the plurality of digital images of H&E stained training slides of training tissue samples is unlabeled.
[0042] In some examples, the single-scale deep learning framework is a convolutional neural network with a ResNet configuration or an Inception-v3 configuration.
[0043] In some examples, the one or more biomarkers are selected from the group consisting of consensus molecular subtypes (CMS) and homologous recombination deficiency ("HRD").
[0044] In some examples, the one or more processors are one or more graphics processing units (GPUs), tensor processing units (TPUs), and / or central processing units (CPUs).
[0045] In some examples, a computing device (e.g., an image-based biomarker prediction system) is communicatively coupled to a pathology slide scanner system via a communications network, whereby the image-based biomarker prediction system receives digital images from the pathology slide scanner system via the communications network.
[0046] In some examples, the computing device is included within a pathology slide scanner system.
[0047] In some examples, the pathology slide scanner system includes an image-based, adversarially trained, and / or predictive model of microsatellite instability (MSI).
[0048] In some examples, generating a report including the digital image and the digital overlay includes generating the digital overlay to include an overlay element that identifies the tumor content of the digital image or the tumor percentage of the digital image.
[0049] According to another example, a computing device configured to identify biomarkers in a digital image of a hematoxylin and eosin (H&E) stained slide of a target tissue includes one or more memories and one or more processors configured to: receive the digital image; perform an image tiling process on the digital image by separating the digital image into a plurality of tile images, each of the plurality of tile images comprising a different portion of the digital image; apply the plurality of tile images to a multi-scale deep learning framework including one or more trained deep learning multi-scale classifier models, each trained to classify a different tissue classification for each tile image, and determine a tissue classification for each of the plurality of tile images using the multi-scale deep learning framework; identify cells in the digital image using the trained cell segmentation model; and identify a predicted presence of one or more biomarkers associated with the digital image from the tissue classification determined for each tile image and from the identified cells in the digital image.
[0050] According to another example, a computing device configured to identify biomarkers in digital images of hematoxylin and eosin (H&E) stained slides of a target tissue, the computing device comprising one or more memories and one or more processors, includes the following steps: receiving a molecular training dataset of a plurality of training tissue samples, the molecular training dataset including RNA transcriptome counts from sequencing of substantially similar samples associated with each training tissue sample; performing a clustering process on the molecular training dataset to identify one or more molecular data subsets, each corresponding to a distinct biomarker; and clustering the one or more molecular data subsets. and one or more processors configured to: receive, for each of the one or more molecular data subsets, a plurality of digital images of H&E-stained training slides of the training tissue samples corresponding to the respective biomarkers for an image-based biomarker prediction system having one or more processors; generate, for each of the one or more molecular data subsets, a trained image-based biomarker classifier model based on the plurality of digital images of the H&E-stained training slides; receive subsequent digital images of H&E-stained slides of the subsequent tissue samples; and apply the subsequent digital images to the trained image-based biomarker classifier model to identify a predicted presence of the one or more biomarkers in the subsequent tissue samples.
[0051] According to another example, a computing device configured to identify biomarkers in a digital image of a hematoxylin and eosin (H&E) stained slide of a target tissue includes one or more memories and one or more processors configured to: receive the digital image into an image-based biomarker prediction system having the one or more processors; separate the digital image into a plurality of tile images, each of the plurality of tile images comprising a different portion of the digital image; apply the plurality of tile images to a deep learning framework including one or more trained biomarker classification models, each trained to classify a different tissue classification; predict a biomarker classification for each of the plurality of tile images using the one or more trained biomarker classification models; determine a predicted presence of one or more biomarkers in the target tissue from the predicted biomarker classification for each of the tile images; and generate a report including the digital image and a digital overlay visualizing the predicted presence of the one or more biomarkers. [Brief explanation of the drawings]
[0052] The patent or application file contains at least one drawing executed in color. Copies of this patent or patent application publication with color drawing(s) will be provided by the U.S. Patent and Trademark Office upon request and payment of the necessary fee.
[0053] The drawings described below illustrate various aspects of the systems and methods disclosed herein. It should be understood that each figure illustrates one example of an aspect of the systems and methods. [Figure 1] FIG. 1 is a block diagram of a schematic diagram of a prediction system having an imaging-based biomarker prediction system, according to an example. [Figure 2] FIG. 1 is a block diagram of a conventional pathologist's cancer diagnosis workflow process. [Figure 3]FIG. 2 is a block diagram of a schematic diagram of a deep learning framework that may be implemented in the system of FIG. 1 , according to an example. [Figure 4] FIG. 1 is a block diagram of a schematic diagram of machine learning data flow, according to an example. [Figure 5] FIG. 4 is a block diagram of a schematic diagram of a deep learning framework that may be implemented in the systems of FIGS. 1 and 3 to form multiple different marker classification models, according to an example. [Figure 6] FIG. 1 is a block diagram of a process for imaging-based biomarker prediction with an example multi-scale configuration. [Figure 7] FIG. 7 is a block diagram of an exemplary process for determining the status of predicted biomarkers, according to an exemplary embodiment of the process of FIG. 6. [Figure 8] FIG. 1 is a block diagram of a process for imaging-based biomarker prediction, with an example single-scale configuration. [Figure 9] FIG. 4 is a block diagram of a process for generating a biomarker prediction report and overlay map that may be performed by the systems of FIGS. 1 and 3, according to an example. [Figure 10A] 10A and 10B show examples of overlay maps generated by the process of FIG. 9, showing a tissue overlay map (FIG. 10A) and a cell overlay map (FIG. 10B), according to one example. [Figure 10B] 10A and 10B show examples of overlay maps generated by the process of FIG. 9, showing a tissue overlay map (FIG. 10A) and a cell overlay map (FIG. 10B), according to one example. [Figure 11] FIG. 1 is a block diagram of a process for preparing digital images of histopathology slides for classification, according to an example. [Figure 12A] 1 illustrates an example of a neural network architecture that may be used for a classification model, according to an example. [Figure 12B] 1 illustrates an example of a neural network architecture that may be used for a classification model, according to an example. [Figure 12C]1 illustrates an example of a neural network architecture that may be used for a classification model, according to an example. [Figure 13] 1 shows a histopathology image showing tile images for classification, according to an example. [Figure 14] FIG. 10 is a block diagram of a schematic diagram of an imaging-based biomarker prediction system using separate pipelines, according to another example. [Figure 15A] FIG. 15 is a block diagram of a schematic diagram of an exemplary biomarker prediction process that may be implemented by the system of FIG. 14, according to an example. [Figure 15B] FIG. 15 is a schematic block diagram of an exemplary training process that may be implemented by the system of FIG. 14, according to an example. [Figure 16] An example input histopathology image is shown. Figures 16A-16C show a representative example of PD-L1 positive biomarker classification. Figure 16A shows the input H&E image, Figure 16B shows the probability map overlaid on the H&E image, and Figure 16C shows PD-L1 IHC staining for reference. Figures 16D-16F show a representative example of PD-L1 negative biomarker classification. Figure 16D shows the input H&E image, Figure 16E shows the probability map overlaid on the H&E image, and Figure 16F shows PD-L1 IHC staining for reference. The color bar indicates the predicted probability of tumor PD-L1+ class. [Figure 17] FIG. 16 is a block diagram of an exemplary multi-view strategy for PD-L1 classification, which may be performed by the process of FIGS. 14, 15A, and 15B, according to an example. [Figure 18] FIG. 1 is a block diagram of a schematic machine learning architecture capable of performing label-free annotation training in a deep learning framework and having a multi-instance learning controller, according to an example. [Figure 19] FIG. 20 is a block diagram of a framework operation that may be implemented by the multi-instance learning controller of FIG. 18, according to an example. [Figure 20]FIG. 20 is a block diagram of a framework operation that may be implemented by the multi-instance learning controller of FIG. 18, according to an example. [Figure 21] FIG. 20 is a block diagram of a framework operation that may be implemented by the multi-instance learning controller of FIG. 18, according to an example. [Figure 22] FIG. 20 is a block diagram of a framework operation that may be implemented by the multi-instance learning controller of FIG. 18, according to an example. [Figure 23] 1 is an example of a resulting overlap map showing biomarker classification of CMS, according to an example. [Figure 24] FIG. 20 is a block diagram of another framework operation that may be implemented by the multi-instance learning controller of FIG. 18, according to another example. [Figure 25] 10 is an example of overlap map results showing biomarker classification for CMS according to another example. [Figure 26] FIG. 20 is a block diagram of another framework operation that may be implemented by the multi-instance learning controller of FIG. 18, according to another example. [Figure 27] According to another example, an example of a neural network architecture that may be used for a classification model is shown. [Figure 28] FIG. 1 is a block diagram of a process for determining a list of potential corresponding therapies (e.g., immunotherapy), according to an example. [Figure 29] FIG. 1 is a block diagram of data flow for generating a list of potential corresponding treatments, according to an example. [Figure 30] FIG. 1 is a block diagram of a system for performing imaging-based biomarker prediction in conjunction with a pathology scanner system, according to an example. [Figure 31] 1, 3, and 30 show various screenshots of an exemplary generated graphic user interface display that may be generated by a system such as those of FIGS. [Figure 32]1, 3, and 30 show various screenshots of an exemplary generated graphic user interface display that may be generated by a system such as those of FIGS. [Figure 33] 1, 3, and 30 show various screenshots of an exemplary generated graphic user interface display that may be generated by a system such as those of FIGS. [Figure 34] 1, 3, and 30 show various screenshots of an exemplary generated graphic user interface display that may be generated by a system such as those of FIGS. [Figure 35] 1, 3, and 30 show various screenshots of an exemplary generated graphic user interface display that may be generated by a system such as those of FIGS. [Figure 36] 1, 3, and 30 show various screenshots of an exemplary generated graphic user interface display that may be generated by a system such as those of FIGS. [Figure 37] 1, 3, and 30 show various screenshots of an exemplary generated graphic user interface display that may be generated by a system such as those of FIGS. [Figure 38] FIG. 1 is a block diagram of an exemplary computing device for use in implementing various systems herein, according to one example. DETAILED DESCRIPTION OF THE INVENTION
[0054] The imaging-based biomarker prediction system comprises a deep learning framework configured and trained to learn directly from pathology tissue slides and predict the presence of biomarkers in medical images. The deep learning framework can be configured and trained to analyze medical images and identify biomarkers that indicate the presence of a tumor, the state / condition of the tumor, or information about the tumor in the tissue sample.
[0055] In an implementation, a cloud-based deep learning framework is used for medical image analysis. Deep learning algorithms automatically learn advanced imaging features to enhance diagnosis, prognosis, treatment indication, and treatment response prediction. In an example, the deep learning framework can directly connect to cloud storage and leverage resources on the cloud platform for efficient deep learning algorithm training, comparison, and deployment.
[0056] In some examples, deep learning frameworks include multi-scale configurations that use tiling strategies to accurately capture the structural and local tissue structure of various diseases (e.g., predicting cancer tumors). These multi-scale configurations perform classification of histopathology images (labeled or unlabeled) using classifiers trained to classify tiles of received histopathology images. In some examples, the multi-scale configurations include tile-level tissue classifiers, i.e., classifiers trained using tile-based deep learning training. In some examples, the multi-scale configurations include pixel-level cell classifiers and cell segmentation models. In some examples, classifications from the tile-level tissue classifiers and pixel-level cell classifiers are analyzed to predict the status of biomarkers in the histopathology images. In further examples, the multi-scale configurations include tile-level biomarker classifiers. Once tracked, the multi-scale classifiers can receive new labeled or unlabeled histopathology images and predict the presence of specific biomarkers in the associated histopathology slides.
[0057] In some examples, the deep learning framework herein includes a single-scale configuration trained using a multi-instance learning (MIL) strategy to predict the presence of biomarkers in histopathology images. A classifier trained using the single-scale configuration can be trained to perform classification of (labeled or unlabeled) histopathology images using a classifier trained using one or more multi-instance learning (MIL) techniques. In some examples, the single-scale configuration includes a slide-level classifier trained using gene sequencing data, such as RNA sequencing data, to analyze histopathology images with slide-level labels rather than tile-level labels. That is, a slide-level classifier is trained using RNA sequencing data to develop an image-based classifier that can predict the status of biomarkers in histopathology images.
[0058] Both the multi-scale and single-scale configurations herein can incorporate various algorithmic optimizations to accelerate computations for such disease analysis.
[0059] In a multi-scale classifier implementation, the deep learning framework can be trained to include classifiers that perform automatic cell segmentation, determine cell / biomarker types, and determine tissue-type classification from histopathology images, thereby providing image-based biomarker development. Even single-scale classifiers can be trained to include tissue-type classification and biomarker classification.
[0060] In the case of a multi-scale classifier configuration, for example, aggregate and spatial imaging features for various cell types (e.g., tumor, stromal, lymphocyte) in digital hematoxylin and eosin (H&E) slides can be determined by a deep learning framework and used to predict clinical and treatment outcomes. Instead of rudimentary manual cell type classification, examples herein use a multi-scale configuration in a deep learning framework to classify each subregion of an H&E slide histopathology image into specific cell segmentations, cell types, and tissue types. From there, biomarker detection is performed by another deep learning framework configured to identify various types of imaging metrics. Examples of imaging metrics include tumor shape, including minimum and maximum tumor shape; tumor area; tumor perimeter; tumor %; cell shape, including cell area; cell perimeter; cell convex area ratio; cell circularity; cell convex perimeter area; cell length; lymphocyte %; cell characteristics; and cell texture, including saturation, intensity, and hue.
[0061] Examples of tissue classes include, but are not limited to, tumor, stroma, normal, lymphocyte, fat, muscle, vascular, immune cluster, necrosis, hyperplasia / dysplasia, erythrocyte, and tissue classes or cell types that are positive (particularly containing an amount of the IHC staining target molecule greater than a certain threshold) or negative (not containing the molecule or containing an amount of the molecule less than a certain threshold) for an IHC staining target molecule.
[0062] In some instances, biomarker detection can be enhanced by combining imaging metrics with structured clinical and sequencing data to develop enhanced biomarkers.
[0063] Biomarkers can be identified by any of the following models: Any model referred to herein may be implemented as an artificial intelligence engine and may include a gradient boosting model, a random forest model, a neural network (NN), a regression model, a naive Bayes model, or a machine learning algorithm (MLA). The MLA or NN can be trained with a training dataset. In an exemplary predictive profile, the training dataset may include imaging, medical condition, clinical, and / or molecular reports, as well as patient details (e.g., curated from an EHR or gene sequencing report). MLAs include supervised algorithms (e.g., algorithms where features / classifications in the dataset are annotated) using linear regression, logistic regression, decision trees, classification and regression trees, naive Bayes, and nearest neighbor clustering; unsupervised algorithms (e.g., algorithms where features / classifications in the dataset are not annotated) using apriori, average clustering, principal component analysis, random forests, and adaptive boosting; and semi-supervised algorithms (e.g., algorithms where an incomplete number of features / classifications in the dataset are annotated) using generative approaches (e.g., Gaussian mixtures, multinomial mixtures, hidden Markov models), sparse separation, graph-based approaches (e.g., minimum cuts, harmonic functions, manifold regularization), heuristic approaches, or support vector machines. NNs include conditional random fields, convolutional neural networks, attention-based neural networks, deep learning, long short-term memory networks, or other neural models. The training dataset includes pathology reports covering multiple tumor samples, RNA expression data for each sample, and imaging data for each sample. Although MLA and neural networks identify different approaches to machine learning, these terms may be used interchangeably herein. Thus, unless otherwise specified, a reference to an MLA may include a corresponding NN, and a reference to a NN may include a corresponding MLA.Training may involve providing an optimized dataset, labeling these characteristics found in patient records, and training the MLA to predict or classify based on new inputs. Artificial neural networks (NNs) are efficient computational models and have demonstrated strengths in solving difficult problems in artificial intelligence. They have been shown to be universal approximators (capable of representing a wide variety of functions given appropriate parameters). Some MLAs can identify important features and their associated coefficients or weights. The coefficients can be multiplied by the frequency of the feature's occurrence to generate a score, and if the score of one or more features exceeds a threshold, the MLA can predict a specific classification. The coefficient schema can be combined with a rule-based schema to generate more complex predictions, such as predictions based on multiple features. For example, 10 key features may be identified for various classifications. A list of key feature coefficients may exist, and a rule set for classification may exist. The rule set may be based on the number of feature occurrences, scaled weights of the features, or other qualitative and quantitative evaluations of the features, coded in logic known to those skilled in the art. In other MLAs, features may be organized in a binary tree structure. For example, a key feature that can distinguish most classifications may exist as the root of a binary tree and each subsequent branch in the tree until a classification is assigned based on reaching a terminal node in the tree. For example, a binary tree may have a root node that tests a first feature. The occurrence or non-occurrence of this feature must exist (a binary decision), and logic can traverse the branch that is true for the item being classified. Additional rules may be based on thresholds, ranges, or other qualitative and quantitative tests. Supervised methods are useful when the training dataset has many known values or annotations, but the nature of EMR / EHR documents may provide fewer annotations. When exploring large amounts of unlabeled data, unsupervised methods are useful for binning / bucketing instances in the dataset.As used herein, a single instance of the above model, or two or more such instances may be combined to constitute a model for purposes of modeling, artificial intelligence, neural networks, or machine learning algorithms.
[0064] In some examples, the technology provides machine learning-assisted histopathology image review, including automatically identifying and outlining tumor regions and / or characteristics of regions or cell types within the regions (e.g., lymphocytes, PD-L1 positive cells, tumors with high tumor budding, etc.), counting cells within the tumor regions, and generating a decision score to improve the efficiency and objectivity of pathology slide review.
[0065] As used herein, the term "biomarker" refers to image-derived information, particularly information in the form of morphological features discernible in histologically stained samples, related to the screening, diagnosis, prognosis, treatment, selection, disease monitoring, progression, and disease recurrence of cancer or other diseases. In some instances, biomarkers herein may be morphological features determined from labeling-based images. Biomarkers herein may be morphological features determined from labeled RNA data.
[0066] A biomarker herein can be image-derived information that correlates with the presence of or susceptibility to cancer in a subject, the likelihood that the cancer is of one subtype or another, the presence or proportion of a biological characteristic such as a type or class of tissue, cell, or protein, the probability that a patient will respond or not respond to a particular treatment or class of treatment, the degree of positive response expected to a treatment or class of treatment (e.g., survival time and / or progression-free survival), whether a patient is responding to a treatment, or the likelihood that the cancer will regress, progress, or progress beyond the site of origin (i.e., metastasize).
[0067] Examples of biomarkers predicted from histopathological images using the various techniques herein include:
[0068] As used herein, tumor-infiltrating lymphocytes (TILs) refer to mononuclear immune cells that infiltrate tumor tissue or stroma. TILs include, for example, T cells, B cells, and NK cells, and their populations can be subdivided based on function, activity, and / or biomarker expression. For example, TIL populations can include, for example, cytotoxic T cells expressing CD3 and / or CD8, and regulatory T cells (also known as suppressor T cells), which are often characterized by FOXP3 expression. Information regarding the density, location, organization, and composition of TILs provides valuable insights into prognosis and potential treatment options. In various embodiments, the present disclosure provides methods for predicting TIL density in a sample, distinguishing subpopulations of TILs in a sample (e.g., distinguishing CD3 / CD8-expressing cytotoxic T cells from FOXP3 Tregs), distinguishing stromal from intratumoral TILs, and the like.
[0069] Programmed death-ligand 1 (PD-L1) is a 40 kDa type 1 transmembrane protein that affects immune system suppression, particularly in patients with autoimmune diseases, cancer, and other pathologies. In the context of cancer immunotherapy, PD-L1 is expressed on the surface of tumor cells, tumor-associated macrophages (TAMs), and T lymphocytes, and can subsequently inhibit PD-1-positive T cells.
[0070] Ploidy refers to the number of sets of homologous chromosomes in a cell's or organism's genome. Examples include haploid, meaning one set of chromosomes, and diploid, meaning two sets of chromosomes. The presence of multiple sets of paired chromosomes in an organism's genome is described as "polyploidy." Three sets of chromosomes, 3n, are triploid, and four sets of chromosomes, 4n, are tetraploid. Extremely large sets can be designated by a number (e.g., 15 sets is 15ploid).
[0071] The nucleus-to-cytoplasm (NC) ratio is a measure of the ratio between the size of a cell's nucleus and the size of that cell's cytoplasm. The NC ratio can be expressed as a volume ratio or cross-sectional area. The NC ratio can indicate cell maturity, as the size of the cell nucleus decreases with cell maturity. In contrast, a high NC ratio in a cell may indicate malignancy of the cell.
[0072] Signet ring morphology is a morphology of signet ring cells, i.e., cells with large vacuoles, that primarily manifests in malignant forms in carcinomas. Signet ring cells are most commonly associated with gastric cancer, but they can arise from a variety of tissues, including the prostate, bladder, gallbladder, breast, colon, ovarian stroma, and testis. For example, signet ring cell carcinoma (SRCC) is a rare form of highly malignant adenocarcinoma. It is an epithelial malignant tumor characterized by the histological appearance of signet ring cells.
[0073] These biomarkers, TIL, NC ratio, ploidy, signet ring morphology, and PD-L1, are examples of biomarkers of morphological features determined from labeling-based images by the techniques herein.
[0074] Consensus molecular subtypes ("CMS") are a set of classification subtypes of colorectal cancer (CRC) developed based on comprehensive gene expression profile analysis. CMS classifications for primary colorectal cancer include CMS1-immune infiltration (often BRAFmut, MSI-High, TMB-High), CMS2-canonical (often ERBB / MYC / WNT-driven), CMS3-metabolic (often KRASmut), and CMS4-mesenchymal (often TGF-B-driven). More broadly, CMS herein includes these and other subtypes of colorectal cancer. Even more broadly, CMS herein refers to subtypes derived from comprehensive gene expression profile analysis of other cancer types described herein.
[0075] Homologous recombination deficiency ("HRD") status is a classification indicating a defect in the normal homologous recombination DNA damage repair process, resulting in loss of replication of a chromosomal region, termed genomic loss of heterozygosity (LOH).
[0076] Biomarkers such as CMS and HRD are examples of biomarkers of morphological features determined from labeled RNA data by the techniques herein.
[0077] By way of example, biomarkers herein include HRD status, DNA ploidy score, karyotype, CMS score, chromosomal instability (CIN) status, signet ring morphology score, NC ratio, activation status of cellular pathways, cell state, tumor characteristics, and splice variants.
[0078] As used herein, a "histopathology image" refers to a digital (including digitized) image of microscopically and histopathologically developed tissue. An example is an image of a histologically stained specimen tissue, where histological staining is a process performed in the preparation of a sample tissue to aid in microscopic examination. In some examples, the histopathology image is a digital image of a hematoxylin and eosin (H&E)-stained histopathology slide, an immunohistochemistry (IHC)-stained slide, a Romanowsky-Giemsa-stained slide, a Gram-stained slide, a Trichrome-stained slide, a Carmine-stained slide, and a silver nitrate-stained slide. Other examples include blood smear slides and tumor smear slides. In other examples, the histopathology image is of other stained slides known in the art. As used herein, references to digital images, digitized images, slide images, and medical images refer to a "histopathology image."
[0079] These histopathology images can also be captured in the visible wavelength range and beyond, such as infrared digital images obtained using spectroscopic examination of tissues with histopathological development. In some examples, the histopathology images include z-stack images representing horizontal cross-sections of a three-dimensional specimen or histopathology slide, captured at various levels of the specimen or at various focal points of the slide. In some examples, the two or more images may be from adjacent or nearly adjacent sections of tissue from the specimen, and one of the two or more images may have a tissue feature that corresponds to a tissue feature in another of the two or more images. There may be a vertical and / or horizontal shift between the location of a corresponding tissue feature in a first image and the location of a corresponding tissue feature in a second image. Therefore, a histopathology image also refers to an image, set of images, or video generated from multiple different images. It should be understood that the following exemplary embodiments can be interchanged or trained with different staining styles, unless expressly excluded.
[0080] Various examples herein are described with reference to a particular class of histopathological images: H&E slide images. Digital H&E slide images can be generated by capturing digital photographs of H&E slides. Alternatively, or in addition, such images can be generated from images derived from unstained tissue via machine learning systems, such as deep learning. For example, digital H&E slide images can be generated from wide-field autofluorescence images of unlabeled tissue sections. See, for example, Rivenson et al., "Virtual histological staining of unlabeled tissue-autofluorescence images via deep learning," Nature Biomedical Engineering, 3(6):466, 2019.
[0081] FIG. 1 illustrates a predictive system 100 that can analyze digital images of histopathology slides of a tissue sample to determine the likely presence of a biomarker in that tissue, where the presence of the biomarker indicates other information about the tumor in the tissue sample, such as the predicted presence of a tumor, the predicted tumor state / condition, or the likelihood of a clinical response to the use of a treatment associated with the biomarker.
[0082] System 100 includes an imaging-based biomarker prediction system 102 that implements, among other things, image processing operations, a deep learning framework, and report generation operations to analyze histopathological images of tissue samples and predict the presence of biomarkers in the tissue samples. In various examples, system 100 is configured to predict the presence of these biomarkers, the tissue locations associated with these biomarkers, and / or the cellular locations of these biomarkers.
[0083] The imaging-based biomarker prediction system 102 can be implemented on one or more computing devices, such as a computer, tablet, or other mobile computing device, or a server, such as a cloud server. The imaging-based biomarker prediction system 102 may include multiple processors, controllers, or other electronic components for processing or facilitating image capture, generation, or storage and image analysis, as well as deep learning tools for analyzing the images, as described herein. An exemplary computing device 3800 for implementing the imaging-based biomarker prediction system 102 is shown in FIG. 38.
[0084] As shown in FIG. 1 , the imaging-based biomarker prediction system 102 is connected to one or more medical data sources via a network 104. The network 104 can be a public network such as the Internet, a private network such as a research or corporate private network, or any combination thereof. The network can include a local area network (LAN), a wide area network (WAN), cellular, satellite, or other network infrastructure, whether wireless or wired. The network 104 can be part of a cloud-based platform. The network 104 can utilize communication protocols including packet-based and / or datagram-based protocols such as Internet Protocol (IP), Transmission Control Protocol (TCP), User Datagram Protocol (UDP), or other types of protocols. Additionally, the network 104 can include multiple devices that facilitate network communication and / or form the hardware infrastructure of the network, such as switches, routers, gateways, access points (such as wireless access points as shown), firewalls, base stations, repeaters, backbone devices, etc.
[0085] The imaging-based biomarker prediction system 102 is communicatively coupled via a network 104 to receive medical images, such as histopathology slides, such as digital H&E-stained slide images, IHC-stained slide images, or digital images of other staining protocols from a variety of different sources. These sources may include a physician's clinical record system 106 and a histopathology imaging system 108. The system 100 can be used to access any number of medical image data sources. Histopathology images may be images captured by any suitable optical histopathology slide scanner, including any dedicated digital medical image scanner, e.g., 20x and 40x resolution magnification scanners. Additionally, the biomarker prediction system 102 can receive images from a histopathology image repository 110. In yet another example, images may be received from a partner genome sequencing system 112, such as TCGA and NCI Genomic Data Commons. Additionally, the biomarker prediction system 102 can receive histopathology images from an organoid modeling lab 116. These image sources can communicate image data, genomic data, patient data, treatment data, historical data, etc., according to the techniques and processes described herein. Each of the image sources can represent multiple image sources. Furthermore, each of these image sources can be considered a different data source, which may generate and provide different image data than other providers, hospitals, etc. The imaging data between the different sources can potentially differ in one or more respects, resulting in biases inherent in the various data sources, such as different dyes, biological sample fixation, embedding, and staining protocols, and different pathology imaging equipment and settings.
[0086] 1 , the imaging-based biomarker prediction system 102 includes an image pre-processing subsystem 114 that performs initial image processing to enhance image data for faster processing in training a machine learning framework and for performing biomarker prediction using the trained deep learning framework. In the illustrated example, the image pre-processing subsystem 114 performs a normalization process on the received image data, including one or more of color normalization 114a, intensity normalization 114b, and imaging source normalization 114c, to compensate for and correct differences in the received image data. In some examples, the imaging-based biomarker prediction system 102 receives medical images, while in other examples, the subsystem 114 can generate medical images from either received histopathology slides or other received images, for example, by aligning shifted histopathology images to compensate for vertical / horizontal shifts to generate a composite histopathology image. This image preprocessing enables deep learning frameworks to more efficiently analyze images across large datasets (e.g., over 1,000, 10,000, 100,000, or 1,000,000 medical images), thereby speeding up the training and analysis process.
[0087] The image pre-processing subsystem 114 may perform further image processing to remove artifacts and other noise from the received images by performing preliminary tissue detection 114d, for example, to identify regions of the image corresponding to histopathologically stained tissue for subsequent analysis, classification, and segmentation.
[0088] As further described herein, in a multi-scale configuration where image data is analyzed on a tile basis, in some examples, image pre-processing includes receiving an initial histopathology image at a first image resolution, downsampling the image to a second image resolution, and then performing normalization on the downsampled histopathology image, such as color and / or intensity normalization, to remove non-tissue objects from the image.
[0089] In contrast, the single-scale configuration does not use downsampling of the received histopathology images and analyzes the image data on a slide-level basis, rather than a tile-based basis.
[0090] In some further hybrid versions of each of the multi-scale and single-scale configurations, a tiling process is applied to the received histopathology images to generate tiles for tile-based analysis.
[0091] The imaging-based biomarker prediction system 102 may be a standalone system that interfaces with external (i.e., third-party) network-accessible systems 106, 108, 110, 112, and 116. In some examples, the imaging-based biomarker prediction system 102 may be integrated with one or more of these systems, including as part of a distributed, cloud-based platform. For example, the system 102 may be integrated with a pathology tissue imaging system, such as a digital H&E stain imaging system, to enable, for example, rapid biomarker analysis and reporting at the imaging station. Indeed, any of the functionality described in the technology herein may be distributed across one or more network-accessible devices, including cloud-based devices.
[0092] In some examples, the imaging-based biomarker prediction system 102 is part of a comprehensive biomarker prediction, patient diagnosis, and patient treatment system. For example, the imaging-based biomarker prediction system 102 can be coupled to communicate predicted biomarker information, tumor prediction, and tumor status information to external systems, including a computer-based pathology lab / oncology system 118, which can receive the generated biomarker report including the image overlay mapping and use it to further diagnose the patient's cancer status and identify corresponding treatments for use in treating the patient. The imaging-based biomarker prediction system 102 can further transmit the generated report to the patient's primary care provider's computer system 120 and a physician's clinical record system 122 for database compilation of the patient report using a database of previously generated reports for the patient and / or generated reports for other patients for use in future patient analyses (including deep learning analyses described herein).
[0093] To analyze the received histopathological imaging data and other data, the imaging-based biomarker prediction system 102 includes a deep learning framework 150 that implements various machine learning techniques to generate a trained classifier model for image-based biomarker analysis from a received set of image data or a set of image data and other patient information. Using the trained classifier model, the deep learning framework 150 is further used to analyze and diagnose the presence of image-based biomarkers in subsequent images collected from the patient. In this manner, images and other data from previously treated and analyzed patients are utilized through the trained model to provide analysis and diagnosis capabilities for future patients.
[0094] In the exemplary system 100, the deep learning framework 150 includes a histopathology image-based classifier training module 160 that can access data received and stored from external systems 106, 108, 110, 112, and 116, as well as any other systems, where the data can be analyzed from received data streams and databased into various data types. The various data types can be divided into image data 162a, which can be associated with other data types: molecular data 162b, demographic data 162c, and tumor response data 162d. The associations can be formed by labeling the image data 162a with one or more different data types. By labeling the image data 162a according to its associations with other data types, the imaging-based biomarker prediction system can train the image classifier module to predict one or more different data types from the image data 162a.
[0095] In the illustrated data, deep learning framework 150 includes image data 162a. For example, to train or use a multi-scale PD-L1 biomarker classifier, this image data 162a may include preprocessed image data received from subsystem 114, images from H&E slides, or images from IHC slides (with or without human annotation), including IHC slides targeting PD-L1, PTEN, EGFR, beta-catenin / catenin beta 1, NTRK, HRD, PIK3CA, and hormone receptors such as HER2, AR, ER, and PR. To train or use other biomarker classifiers, whether multi-scale or single-scale classifiers, image data 162A may also include images from other stained slides. Additionally, in the example of training a single-scale classifier, image data 162A is image data associated with RNA-seq data for a particular biomarker cluster, enabling the multi-instance learning (MIL) techniques herein.
[0096] Molecular data 162b may include DNA sequencing, RNA sequencing, metabolomics data, proteomics / cytokine data, epigenomic data, organoid data, biokaryotype data, transcriptional data, transcriptomics, metabolomics, microbiomics, and immunology, and may include identification of SNPs, MNPs, InDelta, MSI, TMB, CNV fusions, loss of heterozygosity, and loss or gain of function. Epigenomic data includes DNA methylation, histone modifications, or other factors that inactivate genes or cause altered gene function without changing the nucleotide sequence of the gene. Microbiomics includes data on viral infections, which can affect the treatment and diagnosis of certain diseases, and data on bacteria present in a patient's gastrointestinal tract, which can affect the effectiveness of medications taken by the patient. Proteomic data include information about protein composition, structure, and activity, when and where proteins are expressed, rates of protein production, degradation, and steady-state abundance, how proteins are modified (e.g., post-translational modifications such as phosphorylation), protein trafficking between subcellular compartments, protein participation in metabolic pathways, how proteins interact with each other, and post-translational modifications of RNA to proteins such as phosphorylation, ubiquitination, methylation, acetylation, glycosylation, oxidation, or nitrosylation.
[0097] The deep learning framework 150 may further include demographic data 162c and tumor response data 162d (e.g., including data regarding the reduction in tumor growth after exposure to a particular therapy, such as immunotherapy, DNA damaging therapy such as a PARP inhibitor or platinum, or an HDAC inhibitor). The demographic data 162c may include age, sex, race, country of origin, etc. The tumor response data 162d may include epigenomic data, examples of which include changes in chromatin morphology and histone modifications.
[0098] Tumor response data 162d may include cellular pathways, examples of which include IFN gamma, EGFR, MAP kinase, mTOR, CYP, CIMP, and AKT pathways, as well as pathways downstream of HER2 and other hormone receptors. Tumor response data 162d may include cellular state indicators, examples of which include collagen composition, appearance, or refractive index (e.g., extracellular vs. fibroblastic, nodular fasciitis), stromal density or other stromal characteristics (e.g., stromal thickness, wet vs. dry), and / or the general appearance of angiogenesis or vasculature (including the distribution of vasculature in the collagen / stroma, also referred to as epithelial-mesenchymal transition or EMT). Tumor response data 162d may include tumor characteristics, such as tumor complexity, the presence of tumor budding or other morphological features / characteristics indicative of tumor size (including bulky or low-profile tumors), tumor aggressiveness (e.g., known as high-grade basaloid tumors, particularly in colorectal cancer, or high-grade dysplasia, particularly in Barrett's esophagus), and / or the immune status of the tumor (e.g., inflammatory / "hot" tumors versus non-inflammatory / "cold" tumors versus immune-excluded tumors).
[0099] Histopathology image-based classifier training module 160 may be configured with machine learning techniques adapted for image analysis, including, for example, deep learning techniques, such as CNN models, more specifically, tile-resolution CNNs, which in some examples are implemented as FCN models, more specifically, tile-resolution FCN models. Any of data types 162a-162d may be contained within histopathology images and obtained directly from data communicated to imaging-based biomarker prediction system 102, such as data communicated with histopathology images. Data types 162a-162d may be used by histopathology image-based classifier training module 160 to develop classifiers for identifying one or more biomarkers discussed herein.
[0100] In one example, a histopathology image can be segmented, and each segment of the image can be labeled according to one or more data types that can be assigned to that segment. In another example, the histopathology image as a whole can be labeled according to one or more data types that can be assigned to the image or at least one segment of the image. The data types can indicate one or more biomarkers, and labeling the histopathology image or segments with data types can identify the biomarkers.
[0101] In exemplary system 100, deep learning framework 150 further includes a trained image classifier module 170, which may be configured with deep learning techniques, including those that implement module 160. In some examples, trained image classifier module 170 accesses image data 162 for analysis and biomarker classification. In some examples, module 170 further accesses molecular data 162, demographic data 162c, and / or tumor response data 162d for analysis and tumor prediction, corresponding treatment prediction, etc.
[0102] The trained image classifier module 170 includes a trained tissue classifier 172, which has been trained by module 160 using one or more training image sets to identify and classify tissue types in regions / areas of received image data. In some examples, these trained tissue classifiers are trained to identify biomarkers through tissue classification, and include a single-scale configuration classifier 172a and a multi-scale classifier 172b.
[0103] Module 170 may further include other trained classifiers, including a trained cell classifier 174 that identifies biomarkers through cell classification. Module 170 may further include a cell segmenter 176 that identifies cells in the histopathology image, including cell boundaries, cell interiors, and cell exteriors.
[0104] In the examples herein, the tissue classifier 172 may include a biomarker classifier specifically trained to identify tumor infiltration (e.g., ratio of lymphocytes in tumor tissue to all cells in tumor tissue), PD-L1 (e.g., positive or negative status), ploidy (e.g., score), CMS (e.g., subtype discrimination), NC ratio (e.g., nuclear size discrimination), signet ring morphology (e.g., signet cell classification or vacuole size), HRD (e.g., by score or positive or negative classification), etc. according to the biomarkers herein.
[0105] As described in more detail herein, the trained image classifier module 170 and associated classifiers may be configured with machine learning techniques adapted for image analysis, including, for example, deep learning techniques, including, by way of example, CNN models, more specifically, tile-resolution CNNs, which in some instances are implemented as FCN models, more specifically, tile-resolution FCN models, etc.
[0106] System 102 further includes a tumor report generator 180 configured to receive classification data from trained tissue (biomarker) classifier 172, trained cellular (biomarker) classifier 174, and cell segmenter 172, determine tumor metrics of the image data, and generate digital image and statistical data reports, where such output data may be provided to pathology laboratory 118, primary care physician system 120, genome sequencing system 112, tumor board, tumor board electronic software system, or other external computer system for display or consumption in further processes.
[0107] A traditional cancer diagnostic workflow 200 using histopathology images is shown in Figure 2. A biopsy is performed to collect a tissue sample from a patient. In a medical laboratory, digital histopathology images of the tissue sample are generated (202) using known staining techniques, such as H&E or IHC staining, and a digital medical imager (e.g., a slide scanner). These histopathology images are provided to a pathologist who visually analyzes them to identify tumors within the images (204). The pathologist can optionally receive and analyze the patient's genome sequencing data (e.g., DNA-Seq or RNA-Seq data from a genome sequencing lab) (206). The pathologist then diagnoses other characteristics of the type of cancer in the tumor / cancer cells from the visual analysis of the histopathology slide and any genome sequencing data (208) and generates a pathology report (210).
[0108] 3 illustrates an exemplary implementation of the imaging-based biomarker prediction system 102, and more specifically, the deep learning framework 150 in the form of deep learning framework 300. Framework 300 can be communicatively coupled to receive histopathology imaging data and other data (e.g., molecular data, tumor response data, demographic data, etc.) via network 104 from external systems, such as a physician's clinical record system 106, a histopathology imaging system 108, a genome sequencing system 112, a medical image repository 110, and / or the organoid modeling lab 116 of FIG. 1. The organoid modeling lab 116 can collect various types of data, such as the sensitivity of organoids to drugs (e.g., determined by measuring cell death or cell viability after exposure to a drug), single-cell analysis data, or detection of cellular products (including proteins, lipids, and other molecules) indicative of the presence of specific cell populations, including effector data, stimulatory data, regulatory data, inflammatory data, chemoattractant data, as well as organoid image data, any of which can be stored within molecular data 162b.
[0109] The framework 300 includes a pre-processing controller 302, a deep learning framework cell segmentation module 304, a deep learning framework multi-scale classifier module 306, a deep learning framework single-scale classifier module 307, and a deep learning post-processing controller 308.
[0110] To prepare medical images for multi-scale and single-scale deep learning, in one example, the preprocessing controller 302 includes a normalization process 310, which may include color normalization, intensity normalization, and imaging source normalization. The normalization process 310 is optional and can be omitted to facilitate deep learning training, image analysis, and / or biomarker prediction.
[0111] The image discriminator 314 receives the histopathology image normalized by the normalization process 310 and examines the image, including the image metadata, to determine the type of image. The image discriminator 314 can analyze the image data to determine whether the image is a training image, e.g., whether it is an image from a training dataset. The image discriminator 314 can analyze the image data to determine the type of labeling on the image, e.g., whether the image has tile-level labeling, slide-level labeling, or no labeling. The image discriminator 314 can analyze the image data to determine the slide stain used to generate the digital image, H&E, IHC, etc.
[0112] In response to examining this image data, the image discriminator 314 determines which images are provided to a slide-level label pipeline 313 to feed into a single-scale classifier module 307 of the deep learning framework, as well as which images are provided to a tile-level label pipeline 315 to feed into a multi-scale classifier 306 of the deep learning framework.
[0113] In the illustrated example, images with tile-level labeling in pipeline 315 include a tissue detection process and an image tiling process. These processes may be performed on all received image data, only on training image data, only on image data received for analysis, or some combination thereof. In some examples, for example, the tissue detection process may be omitted to facilitate deep learning training, image analysis, and / or biomarker prediction. Indeed, any of the controller 302's processes may be executed in a dedicated biomarker prediction system or may be distributed for performance by an externally connected system. For example, a tissue pathology imaging system may be configured to perform a normalization process before sending image data to the biomarker prediction system. In some examples, the biomarker prediction system may communicate an executable normalization software package to connected external systems, which configure their systems to perform normalization or other preprocessing.
[0114] In examples where the image discriminator 314 sends unlabeled images to the pipeline 315, the pipeline 315 includes a multi-instance learning (MIL) controller, described further herein, configured to convert all or a portion of these histopathology images into tile-labeled images. The MIL controller can be configured to perform processes such as those described herein in FIGS. 18-26.
[0115] To facilitate tissue detection in the trained tissue classifier, the tissue detection process of pipeline 315 can perform initial tissue identification to locate and segment tissue regions of interest for biomarker analysis. Such tissue identification of interest may include, for example, identifying tissue boundaries and segmenting the image into tissue and non-tissue regions, such that metadata identifying the tissue regions is stored with the image data to facilitate processing and prevent attempted biomarker analysis in non-tissue regions or regions that do not correspond to the tissue of interest.
[0116] To facilitate deep learning classification in various multi-scale configurations, the deep learning framework multi-scale classifier module 306 is configured to classify tissue using tiling analysis. For example, in pipeline 315, the tissue detection process sends a pathology tissue image (e.g., image data enriched with tissue detection metadata) to the image tiling process, which selects and applies a tiling mask to the received image to break the image into smaller sub-images for analysis by the framework module 306. The pipeline 315 can store multiple different tiling masks and select a tiling mask. In some examples, the image tiling process selects one or more tiling masks optimized for different biomarkers. That is, in some examples, the image tiling is biomarker-specific. This allows for the use of tiles of various pixel sizes and shapes specifically selected to, for example, improve accuracy and / or reduce processing time associated with a particular biomarker. For example, a tile size optimized for identifying the presence of TILs in an image may be different from a tile size optimized for identifying PD-L1 or another biomarker. Thus, in some examples, the pre-processor controller 302 is configured to perform image processing and tiling specific to a type of biomarker, and after the system 300 analyzes the image data for that biomarker, the controller 302 can reprocess the original image data for analysis for the next biomarker until all biomarkers have been examined.
[0117] Generally speaking, the tiling mask applied by the image tiling process of the pipeline 315 may be selected to increase the efficiency of operation of the deep learning framework module 306. The tiling mask may be selected based on the size of the received image data, based on the configuration of the deep learning framework 306, based on the configuration of the framework module 304, or some combination thereof.
[0118] The tiling masks may have tiling block sizes that vary. Some tiling masks have uniform (i.e., each the same size) tiling blocks. Some tiling masks have tiling block sizes that vary. The tiling mask applied by the image tiling process may be selected based on, for example, the number of classification layers in the deep learning framework 306. In some examples, the tiling mask may be selected based on the processor configuration of the biomarker prediction system, for example, if multiple parallel processors are available or if a graphical processing unit or a tensor processing unit is used.
[0119] In the illustrated example, the deep learning multi-scale classifier module 304 is configured to perform cell segmentation via the cell segmentation model 316. Here, cell segmentation can be a pixel-level process of the histopathology image from the normalization process 310. In other examples, this pixel-level process may be performed on image tiles received from the pipeline 315. In some examples, the cell segmentation process of the framework 304 results in biomarker classification because some of the biomarkers identified herein are determined from a cell-level analysis, as opposed to a tissue-level analysis. These include, for example, signet rings, large nuclei, and high NC ratios. The module 304 can be configured using a CNN configuration, and in particular, an FCN configuration, to implement each separate segmentation.
[0120] The deep learning framework multi-scale classifier module 306 includes a tissue segmentation model 318, a tissue classification model 320, and a biomarker classification model 320. Like module 304, module 306 can be configured using a CNN configuration, and in particular, an FCN configuration, to implement each separate segmentation.
[0121] In one example, the cell segmentation model 316 of module 304 can be configured as a three-class semantic segmentation FCN model developed by modifying a UNet classifier to replace the loss function with a cross-entropy function, a focal loss function, or a mean squared error function to form a three-class segmentation model. The three-class nature of the FCN model means that the cell segmentation model 316 can be configured as a first pixel-level FCN model that identifies and assigns each pixel of the image data to a cell subunit class: (i) cell interior, (ii) cell boundary, or (iii) cell exterior. This is provided as an example. The segmentation size of the module model 316 can be determined based on the type of cell being segmented. For example, for both TIL biomarkers, the model 316 can be configured to perform lymphocyte identification and segmentation using a three-class FCN model. For example, the cell segmentation model 316 can be configured to classify pixels in an image as corresponding to (i) the interior, (ii) the boundary, or (iii) the exterior of a lymphocyte cell. The cell segmentation model 316 can be configured to identify and segment any number of cells, including, for example, tumor positive, tumor negative, lymphocyte positive, lymphocyte negative, immune cells including lymphocytes, cytotoxic T cells, B cells, NK cells, macrophages, etc.
[0122] In some examples, module 304 receives tiled sub-images from pipeline 315, and cell segmentation model 316 determines a list of all lymphocyte locations, which are compared to lists of all cells from the other three class models determined from model 316 to eliminate falsely detected lymphocytes that are not cells. System 300 then obtains the new list of confirmed lymphocyte locations from this module 304 and compares it with the list of tissues from tissue segmenter module 318, e.g., tumor and non-tumor tissue locations determined from tissue classification model 320, to determine whether the lymphocytes are in tumor or non-tumor regions.
[0123] The three-class model facilitates counting individual cells and allows for more accurate classification, especially when two or more cells overlap each other. Tumor-infiltrating lymphocytes overlap with tumor cells. In a traditional two-class cell contour model, which labels pixels only whether they contain the outer edge of a cell, each cluster of two or more overlapping cells is counted as a single cell.
[0124] In addition to using a three-class model, the cell segmentation model 316 can be configured to avoid the possibility of cells that span two tiles being counted twice by adding a buffer around each tile on all four sides that is slightly wider than the average cell. The intention is to count only cells that appear in the central, unbuffered region of each tile. In this case, tiles are positioned so that the central, unbuffered regions of adjacent tiles are adjacent and do not overlap. Adjacent tiles overlap in their respective buffer regions.
[0125] In one example, the cell segmentation algorithm of model 316 can be formed from two UNet models. One UNet model can be trained using images of mixed tissue classes in which a human analyst has highlighted the outer edge of each cell and classified each cell according to its tissue class. In one example, the training data includes digital slide images in which all pixels are labeled as either the interior of a cell, the outer edge of a cell, or the background, which is outside all cells. In another example, the training data includes digital slide images in which all pixels are labeled "yes" or "no" to indicate whether they represent the outer edge of a cell. This UNet model can recognize the outer edges of many types of cells and classify each cell according to its shape or its location within the tissue class region assigned by tissue classification module 320.
[0126] Another UNet model can be trained on images of many cells of a single tissue class, or on images of a diverse set of cells where only cells of one tissue class are outlined with a binary mask. In one example, the training set is labeled by associating a first value with all pixels that represent the cell type of interest and a second value with all other pixels. Visually, an image labeled in this way appears as a black-and-white image, with all pixels that represent the tissue class of interest being white and all other pixels being black, and vice versa. For example, an image might have only lymphocytes labeled. This UNet model can recognize the outer edges of that particular cell type and assign a label to cells of that type within the digital image of the slide.
[0127] The cell segmentation model 316 is a trained cell segmentation model that can be used for cell detection, although in some examples, the model 316 is configured as a biomarker detection model configured as a pixel-level classifier that classifies pixels as corresponding to a biomarker.
[0128] Turning to the multi-scale classifier module 306 of the deep learning framework, the tissue segmentation model 318 may be configured in a similar manner to the segmentation model 316, i.e., as a three-class semantic segmentation FCN model developed by modifying the UNet classifier by replacing the loss function with a cross-entropy function, a focal loss function, or a mean squared error function to form a three-class segmentation model. The model 318 may identify the interior, exterior, and boundary of various tissue types within a tile.
[0129] The tissue classification model 320 is a tile-based classifier configured to classify tiles as corresponding to one of a plurality of different tissue classifications. Examples of tissue classes include, but are not limited to, tumor, stroma, normal, lymphocyte, fat, muscle, vascular, immune cluster, necrosis, hyperplasia / dysplasia, erythrocyte, and tissue classes or cell types that are positive (particularly containing an amount of the IHC staining target molecule greater than a certain threshold) or negative (not containing the molecule or containing an amount of the molecule less than a certain threshold) for an IHC staining target molecule. Examples also include tumor-positive, tumor-negative, lymphocyte-positive, and lymphocyte-negative.
[0130] With the cell segmentation in the histopathology image generated by cell segmentation model 316 and the tissue classification from tissue classification model 302, biomarker classification model 322 receives data from both and determines the presence of predictive biomarkers in the histopathology image, and particularly in a multi-scale configuration, determines the presence of predictive biomarkers in each tile image of the histopathology image. Biomarker classification model 322 can be implemented in deep learning framework model 306 as shown, or can be a trained classifier implemented separately from model 306, such as in deep learning post-processing controller 308.
[0131] In some examples of biomarker classification models for detecting TIL biomarkers, a tissue classification model 320 is trained to identify the percentage of TILs within a tile image, a cell segmenter 316 determines cell boundaries, and a biomarker classification model 322 classifies the tile image based on the percentage of TILs within the cell interior, resulting in a classification of (i) tumor-IHC / lymphocyte positive or (ii) non-tumor-IHC / lymphocyte positive.
[0132] In some examples of biomarker classification models for detecting polyploidy, the biomarker classification model 322 can be trained with a ploidy model based on histopathology images and associated ploidy scores, for example, using techniques provided in Coudray N, Ocampo PS, Sakellaropoulos T, Narula N, Snuderl M, Fenyo D, et al., "Classification and mutation prediction from non-small cell lung cancer histopathology images using deep learning," Nat Med. 2018;24:1559-67.
[0133] In one example, training data can be data such as an actual karyotype determined by a cytogeneticist, but in some examples, the biomarker classification model 322 can be configured to infer such data. Ploidy data can be formatted as a column of chromosome number, start position, stop position, and region length. Ploidy scores can be determined from DNA sequencing data and can be specific to genes, chromosomes, or chromosome arms. Ploidy scores can be global and represent the entire genome of a sample (global CNVs / copy number variations can cause changes in hematoxylin staining of tumor nuclei), or they can be calculated by averaging the ploidy scores of each region within the genome, with local regional ploidy scores weighted according to the length of each region associated with that score. The trained ploidy models of the biomarker classification model 322 can be specific to genes, chromosome arms, or entire chromosomes, as each section can have a different impact on the cellular morphology seen in histopathology images. The predicted ploidy biomarker data can influence the accept / reject analysis, as even if the tumor purity or cell count on a slide is low, if the ploidy is higher than normal there may still be enough material remaining for genetic testing. The biomarker metric processor 326 can be configured to make such a determination prior to report generation.
[0134] In some examples of biomarker classification models for detecting signet ring morphology, the biomarker classification model 322 may be trained with a signet ring morphology model based on classification techniques such as poorly cohesive (PC), signet ring cell (SRC), and Laurent's subclassifications, as well as Mariette, C., Carneiro, F., Grabsch, H.I. et al., "Consensus on the pathological definition and classification of poorly cohesive gastric carcinoma," Gastric Cancer 22, 1-9 (2019) and other signet ring morphology classifications.
[0135] In some examples of biomarker classification models for detecting NC ratios, the cell segmentation model 316 may be configured with the three-class UNet described herein, which is trained to identify three classes: nucleus, cytoplasm, and cell boundary / non-cellular background. In one example, the training data may be images in which each pixel has been manually annotated with one of these three classes and / or may be images so annotated by a trained model, as illustrated in the example of updated training images in FIG. 4.
[0136] Thus, the cell segmentation model 316 can be trained to analyze an input image and assign one of three classes to each pixel, defining a cell as the group of all cytoplasmic pixels between an adjacent nuclear pixel and the next-closest boundary pixel. The biomarker classification model 322 can then be configured to calculate, for each cell, the nuclear:cytoplasmic ratio as the area (in pixels) of the cell's nucleus divided by the area (in pixels) of the entire cell (nucleus and cytoplasm).
[0137] To identify tumor tissue and the tumor status of the tissue, the deep learning framework 306 can be configured, in one example, using an FCN classifier. In one example, the deep learning framework 304 can be configured as a pixel-resolution FCN classifier, while the deep learning framework 306 can be configured as a tile-resolution FCN classification model or a tile-resolution CNN model, i.e., a model that performs classification on an entire received tile of image data.
[0138] The classification model 320 of module 306 can be configured to classify tissue within a tile as corresponding to one of a plurality of tissue classes, such as, for example, biomarker status, tumor status, tissue type, and / or tumor state / condition, or other information. In the illustrated example, module 306 is configured with tissue classification 320 and tissue segmentation model 322. In an exemplary implementation of TIL biomarkers, the tissue classification model 320 can classify tissue using tissue classifications such as tumor-IHC positive, tumor-IHC negative, necrotic, stromal, epithelial, or blood. The tissue segmentation model 328 identifies boundaries of the different tissue types identified by the tissue classification model 320 and generates metadata for use by the post-processing controller 308 to visually display the boundaries and color coding of the different tissue types in an overlay mapping report generator.
[0139] In an example implementation, deep learning framework 300 performs tile-based classification by receiving tiles (i.e., sub-images) from the classification models 320 and 322 processes. In some examples, tiling may be performed by framework 306 using a tiling mask, and module 306 itself may send the generated sub-images to module 304 for pixel-level segmentation in addition to performing tissue classification. Module 306 may examine each tile sequentially, or module 306 may examine each tile in parallel for faster processing of images due to the nature of the matrix generated by the FCN model.
[0140] In some examples, the tissue segmentation model 318 receives pixel-resolution cell segmentation data and / or pixel-resolution biomarker segmentation data from module 304 and performs statistical analysis on a tile-by-tile basis. In some examples, the statistical analysis determines (i) the area of the image data covered by tissue, e.g., the area of a stained histopathology slide covered by tissue, and (ii) the number of cells in the image data, e.g., the number of cells in a stained histopathology slide. For example, the tissue segmentation model 318 can accumulate cell and tissue classifications for each tile of the image until all tiles forming the image have been classified.
[0141] If the deep learning framework is a multi-scale classifier module for tile-based biomarker classification, the deep learning framework 300 is further configured to classify biomarkers using classifications trained from slide-level training images without requiring tile-level labeling. For example, as discussed further below, slide-level training images received by the image discriminator can be provided to a slide-level labeler 313 having an MIL controller configured to perform processes such as those described in Figures 18-26 herein to generate multiple tile images with inferred classifications and, optionally, perform tile selection on the tiles to train a tissue classification model 317 and a biomarker classification model 319. Examples of single-scale classifiers include a CMS biomarker classification model whose output is a CMS class and an HRD biomarker classification model whose output is HRD+ or HRD-. These classifications may be performed on the entire histopathology image to determine biomarker predictions, or on each tile image of the digital image and analyzed by a biomarker metric processor 326 to determine biomarker predictions from the tile images.
[0142] In some examples of biomarker classification models for detecting HRD, the biomarker classification model 319 can be configured to predict HRD. Training of the HRD model in the classifier 318 can be based on histopathology images and matched HRD scores. For example, training data can be generated by H&E images and RNA-seq data, which in some examples includes RNA expression profile data, which is fed back to the HRD model for further training, similar to the updated training data 403 in FIG. 4. The training data can be derived from tumor organoids, pairing H&E images of the organoids with measurements of the organoids' sensitivity to PARP inhibitors, which are indicative of HRD, or results of an HRD model run by the RNA team based on the organoids' RNA expression profiles. Exemplary HRD prediction models of biomarker classification model 319 are described in Peng, Guang et al., "Genome-wide transcriptome profiling of homologous recombination DNA repair," Nature communications vol. 5 (2014): 3361 and van Laar, RK, Ma, X.-J., de Jong, D., Wehkamp, D., Floore, AN, Warmoes, MO, Simon, I., Wang, W., Erlander, M., van't Veer, LJ and Glas, AM (2009), "Implementation of a novel microarray-based diagnostic test for cancer of unknown primary," Int. J. Cancer, 125: 1390-1397.
[0143] In one example, a deep learning framework can identify HRD from H&E slides using RNA expression and identify slide-level labels indicating the percentage of the slide containing biomarker-expressing cells. In one example, an RNA label activation map approach can be applied to the entire slide as a binary label (i.e., positive or negative HRD expression anywhere in the tissue) or as a continuous percentage (i.e., if 62% of the cells in the image are found to express HRD). Binary RNA labels can be generated by next-generation sequencing of the specimen, and cell percentage labels can be generated by applying single-cell RNA sequencing. In one example, single-cell sequencing can identify the cell types and quantities present in RNA expression from NGS.
[0144] Training a tile-based deep learning network to predict biomarker classification labels for each tile of a whole slide image can be performed using any of the methods described herein. Once training is complete, the model can be applied to each tile using activation mapping methods. Activation mapping can be performed using Gradient Class Activation Mapping (Grad-CAM) or guided backpropagation. Both can identify which regions of the tile contribute most to the classification. In one example, the portion of the tile that contributes most to the HRD-positive class can be the cells clustered in the upper right corner of the tile. Cells within the identified active region can then be labeled as HRD-positive cells.
[0145] Proving that a model will perform reliably clinically may involve comparing the model's results to a source of ground truth. One possible method for generating ground truth involves isolating small regions of tissue containing fewer than 100 cells each by segmenting through tissue microarrays and sequencing each region individually to obtain RNA labels for each region. A further method for generating ground truth may involve classifying these regions using a biomarker classification model to determine the accuracy with which the activation map highlights cells in regions of high HRD expression and ignores most cells in regions of low HRD expression.
[0146] When training a tile-based deep learning network to predict biomarker classification labels for each tile, a powerful supervised approach is utilized to generate biomarker labels and identify the HRD status (positive or negative) of individual cells. Single-cell RNA sequencing can be used alone or in combination with laser-guided microdissection to extract one cell at a time and create a label for each cell. In one example, a cell segmentation model can be incorporated to first obtain cell outlines, and then an artificial intelligence engine can be incorporated to classify pixel values within each cell outline according to biomarker status. In another example, a mask of the image can be generated, with HRD-positive cells assigned a first value and HRD-negative cells assigned a second value. The slide with the mask can then be used to train a single-scale deep learning framework to identify cells expressing HRD.
[0147] In some examples of biomarker classification models for detecting CMS, the biomarker classification model 319 can be configured to predict CMS. Such biomarker classifications can be configured to classify segmented cells as corresponding to cancer-specific classifications. For example, four trained CMS classifications for primary colorectal cancer include: 1-immune infiltration (often BRAFmut, MSI-High, TMB-High), 2-canonical (often ERBB / MYC / WNT-driven), 3-metabolic (often KRASmut), and 4-mesenchymal (often TGF-B-driven). In other examples, more trained CMS classifications can be used, but generally, two or more CMS subtypes are classified in the examples herein. Additionally, other cancer types may have their own trained CMS categories, and the classifier 318 can be configured with models for subtyping each cancer type. Examples of techniques for developing CMS classifications for 4, 5, 6, 7 or more CMS classifications are described in Eide, PW, Bruun, J., Lothe, RA et al., "CMScaller: an R package for consensus molecular subtyping of colorectal cancer pre-clinical models," Sci Rep 7, 16618 (2017) and https: / / github.com / peterawe / CMScaller.
[0148] Training of the CMS model in classifier 318 can be based on histopathological images matching CMS category assignments. CMS category assignments can be based on RNA expression profiles, and in one example are generated by an R program called CMS Caller, which uses closest template predictions (see Eide, PW, Bruun, J., Lothe, RA et al., "CMScaller: an R package for consensus molecular subtyping of colorectal cancer preclinical models," Sci Rep 7, 16618 (2017) and https: / / github.com / peterawe / CMScaller). An alternative classification using a random forest model is described in Guinney, J., Dienstmann, R., Wang, X. et al., "The consensus molecular subtypes of colorectal cancer," Nat Med 21, 1350-1356 (2015)). For example, a CMS caller examines each RNA-seq data sample and determines whether each gene is above or below average, resulting in a binary classification of each gene. This avoids batch effects, for example, between different RNA-seq datasets. Training data can also include DNA data, IHC data, mucin markers, and treatment response / survival data from clinical reports. These may or may not be associated with CMS category assignment. For example, CMS4 IHC stains positive for TGFbeta, CMS1 IHC may be positive for CD3 / CD8, CMS2 and 3 have alterations in mucin genes, CMS2 responds to cetuximab, and CMS1 responds well to Avastin. CMS1 has the best survival prognosis, while CMS4 has the worst prognosis. (See slide 12 of the CMS slides.) CMS categories 1 and 4 can be detected from H&E. Using the architecture shown in Figure 4, for example, a model can be trained to distinguish and classify CMS2 from CMS3.
[0149] In one example, a biomarker classification model 319 can be configured to identify five CRC-specific subtypes (CRIS) with unique molecular, functional, and phenotypic specificities: (i) CRIS-A: mucinous, glycolytic, enriched microsatellite instability, or KRAS mutation; (ii) CRIS-B: TGF-β pathway activity, epithelial-to-mesenchymal transition, poor prognosis; (iii) CRIS-C: elevated EGFR signaling, sensitivity to EGFR inhibitors; (iv) CRIS-D: WNT activation, IGF2 gene overexpression and amplification; and (v) CRIS-E: Paneth cell-like phenotype, TP53 mutation. While CRIS subtypes successfully classify independent sets of primary and metastatic CRCs, they have limited overlap with existing transcriptional classes, and their predictive and prognostic performance has been unprecedented. See, for example, Isella, C., Brundu, F., Bellomo, S. et al., "Selective analysis of cancer-cell intrinsic transcriptional traits defines novel clinically relevant subtypes of colorectal cancer," Nat Commun 8, 15107 (2017).
[0150] For biomarker detection, the biomarker classification model 319 may be trained with a CMS model that predicts for each tile, a CMS classifier that identifies different tissue types (e.g., stroma) for that classification, instead of simply attempting an average CMS classification across all tiles. In one example, each tile is processed, and the CMS model generates a condensed representation using the pixel data associated with each tile, and each tile is assigned to a class (cluster 1, cluster 2, etc.) based on the pattern of pixel data for each tile and the similarity between tiles. A list of the percentage of tiles in the image that belong to each cluster is the image's cluster profile, which may be provided by a report generator. In one example, each profile is fed to the model for training along with its corresponding CMS designation or RNA expression profile (which was the original method used to define CMS categories). In another example, each tile in all training slide images is annotated according to the overall CMS category assigned to the entire slide from which the tile originated, and the tiles are clustered and analyzed to determine the cluster most closely associated with a CMS category.
[0151] In some examples, instead of performing the same classification on each tile and weighting the tile classifications equally, the biomarker classification model 319 (and biomarker classification model 322) can cluster each tile into a discrete number of clusters. One way to achieve this is to include an attention layer in the biomarker classification model. In one example, all tiles in all training slides can be classified into clusters, and then if a number of tiles in a cluster are not statistically associated with a biomarker, that cluster will not be weighted as highly as a cluster that is associated with a biomarker. In other examples, a majority voting technique can be used to train the biomarker classification model 319 (or model 322).
[0152] Although shown as separate models, biomarker classification models 322 and 319 may each be configured to include all or a portion of a corresponding tissue classification model, cell segmentation model, and tissue segmentation model, as in the case of classifying various biomarkers herein. Furthermore, while biomarker classification model 322 is shown to be included within multi-scale classifier module 306 and biomarker classification model 319 is shown to be included within single-scale classifier module 307, in some examples, all or a portion of these biomarker classification models may be implemented in post-processing controller 308, as in the case of classifying various biomarkers herein. Furthermore, while described as tile-level or slide-level classification models, in some examples, biomarker classification models 322 and 319 may in some examples be configured as pixel-level classifiers.
[0153] From the decisions made by modules 304, 306, and 307, post-processing controller 308 can determine whether the image data contains an amount of tissue that exceeds a threshold and / or meets a criteria, for example, whether the image data contains enough tissue for genetic analysis, enough tissue to use the image data as training images in the learning phase of a deep learning framework, or enough tissue to combine the image data with an existing trained classifier model.
[0154] Thus, in various examples herein, including those described with reference to FIG. 3 and elsewhere herein, a patient report may be generated that may be presented to a patient, physician, medical professional, or researcher in the form of a digital copy (e.g., a JSON object, a PDF file, or an image on a website or portal), a hard copy (e.g., a printout on paper or another tangible medium), audio (e.g., recorded or streamed), or another format.
[0155] The report may include information related to the gene expression call (e.g., over- or under-expression of a particular gene), the detected genetic variants, other characteristics of the patient's sample, and / or clinical records. The report may further include clinical trials for which the patient is eligible, treatments for which the patient may be eligible, and / or predicted side effects if the patient receives a particular treatment, based on the detected genetic variants, other characteristics of the sample, and / or clinical records.
[0156] The results contained in the report and / or additional results (e.g., from a bioinformatics pipeline) can be used to analyze a database of clinical data, particularly to determine whether there are trends indicating that a treatment will slow the progression of cancer in other patients with the same or similar outcomes as the specimen. The results can also be used to design tumor organoid experiments. For example, organoids can be genetically engineered to have the same characteristics as the specimen and then subjected to a treatment and observed to determine whether the treatment can slow the growth rate of the organoids, and therefore the growth rate of the patient associated with the specimen.
[0157] In one example, the post-processing controller 308 is further configured to determine a plurality of different biomarker predictive metrics and a plurality of tumor predictive metrics, for example, using the biomarker metric processing module 326. Example predictive metrics include tumor purity, number of tiles classified as a particular tissue class, number of cells, number of tumor-infiltrating lymphocytes, cell type or tissue class clustering, cell type or tissue class density, tumor cell characteristics (circularity, length, nuclear density), stromal thickness surrounding the tumor tissue, image pixel data statistics, predicted patient survival, PD-L1 status, MSI, TMB, tumor origin, and immunotherapy / treatment response.
[0158] For example, the biomarker metric processing module 326 can determine, for each tissue class, the number of tiles classified into one or more single tissue classes, the percentage of tiles classified into each tissue class, for any two classes, the ratio of the number of tiles classified into the first tissue class to the number of tiles classified into the second tissue class, and / or the total area of tiles classified into a single tissue class. The module 326 can determine tumor purity based on the number of tiles classified as tumor relative to other tissue classes, or the number of cells located in tumor tiles relative to the number of cells located in other tissue class tiles. The module 326 can determine cell counts for the entire histopathology image, within a predefined area by the user, within tiles classified as one of the tissue classes, within a single grid tile, or across an entire area or region of interest, whether predetermined, selected by the user during operation of the system 300, or automatically selected by the system 300, for example, by determining the most likely region of interest based on image analysis. The module 326 can determine the clustering of cell types in tissue classes based on the spacing and density of classified cells, the spacing and distance of tiles classified into tissue classes, or any visually detectable feature. In some examples, module 326 determines the probability that two adjacent cells are, for example, two immune cells, two tumor cells, or one of each. Module 326 determines tumor cell characteristics by determining the average circularity, perimeter, and / or nuclear density of identified tumor cells. The thickness of the identified stroma can be used as a predictor of patient response to treatment. Image pixel data statistics determined by module 326 may include the mean, standard deviation, and sum of each tile of either a single image or a collection of images of any pixel data, including red, green, and blue (RGB) values, optical density, hue, saturation, grayscale, and stain deconvolution.Additionally, module 326 may calculate line locations, alternating intensity patterns, shape contours, segmented tissue classes, and / or staining patterns of segmented cells within the image. In any of these examples, module 326 may be configured to create a determined / predicted status, and then overlay display generation module 324 generates a report to display the determined information. For example, overlay map generation module 324 may generate a network-accessible user interface that allows a user to select different types of data to be displayed. Module 324 generates an overlay map showing the selected different types of data overlaid on a rendition of the original stained image data.
[0159] FIG. 4 shows a machine learning data input / flow schematic 400 that may be implemented in the system 300 of FIG. 3, or more generally, in any of the systems and processes described herein.
[0160] In training mode, in which the deep learning framework of system 300 is trained, various training data can be acquired. In the illustrated example, training image data 401 in the form of high-resolution and low-resolution histopathology images is provided to preprocessing controller 302. As shown, the training images can include annotated tissue image data from various tissue types, e.g., tumor, stroma, normal, immune cluster, necrosis, hyperplasia / dysplasia, and red blood cells. As shown, the training images can include computer-generated synthetic image data, as well as image data of segmented cells (cellular image data) and image data of biomarkers (e.g., biomarkers discussed herein) labeled with slide-level or tile-level labels (collectively, biomarker-labeled image data). These training images can be digitally annotated, although in some examples, the tissue annotation is performed manually. In some images, the training image data includes molecular data and / or demographic data, e.g., as metadata within the image data. In the illustrated example, such data is separately fed to deep learning framework 402 (consisting of exemplary implementations of multi-scale deep learning framework 306′ and single-scale deep learning framework 307′). Other training data, such as pathway activation scores, may also be provided to controller 302 for additional training of the deep learning framework.
[0161] In some examples, the deep learning framework 402 generates updated training images 403, which are annotated and segmented by the deep learning framework 402 and fed back to the framework 402 (or preprocessing controller 302) for use in updated training of the framework 402.
[0162] In diagnostic mode, patient image data 405 is provided to controller 302 and used in accordance with the examples herein.
[0163] Any of the image data herein, including patient image data and training images, may be histopathological image data, such as H&E slide images and / or IHC slide images. For example, in the case of IHC training images, the images may be segmented images that distinguish between cytotoxic T cells and regulatory T cells, or other cell types.
[0164] In some examples, the controller 302 generates image tiles 407 and accesses one or more tiling masks 409 and tile metadata 411, which are supplied as input to a deep learning framework 402 for the controller 302 to determine predicted biomarkers and / or tumor status and metrics, which are then provided to an overlay report generator 404 for generating a biomarker and tumor report 406. Optionally, the report 406 may include an overlay of the histopathology image and may further include biomarker scoring data, such as percentage TILs, in one example.
[0165] In some examples, clinical data 413 is provided to the deep learning framework 402 for use in analyzing the image data. The clinical data 413 may include health records, biopsy tissue type, and anatomical location of the biopsy. In some examples, tumor response data 415 collected from the patient after treatment is additionally provided to the deep learning framework 402 to determine biomarker status, tumor status, and / or changes in those metrics.
[0166] FIG. 5 illustrates an exemplary deep learning framework 500 formed from multiple different biomarker classification models. The elements of FIG. 5 are provided as follows: "Cell" refers to a cell segmentation model, e.g., a trained pixel-level segmentation model, according to examples herein. "Multi" refers to a multi-scale (tile-based) tissue classification model, according to examples herein. "Post" refers to arithmetic calculations that may be performed in the final stage of a biomarker classification model configured to predict the biomarker status of an image or tile image, according to examples herein, in response to one or more data from the "Cell" or "Multi" stages. In one example, "Post" may include majority voting, such as summing the number of tiles in the image associated with each biomarker label and identifying assigning the biomarker status in the image to the biomarker label with the largest sum. A two-tier "Post" configuration refers to a two-stage post-processing configuration, where arithmetic calculations may be stacked. In one example, the first tier of Post may include summing the tissue and cells in a labeled tile and summing lymphocyte cells in the same tile. The second layer can divide the lymphocyte cell count by the cell count to generate a ratio, which can be used to assign the status of the biomarker in the image based on whether the ratio exceeds the threshold when compared to the threshold. The final “post” configuration can include other post-processing functions, such as the report generation process described herein. “Single” refers to a single-scale classification model, according to examples herein. “MIL” refers to an MIL controller, according to examples herein. In the illustrated example, the deep learning framework 500 includes a TIL biomarker classification model 502, a PD-L1 classification model 504, a first CMS classification model 506 based on a “single” classification architecture, a second CMS classification model 508 based on a “multiple” classification architecture, and an HRD classification model 510. Patient data 512, such as molecular data, demographic data, tumor response data, and patient images 514, are stored in datasets accessible by the deep learning framework 500.
[0167] The training data is also shown in the form of cell segmentation training data 516, single-scale classification biomarker training data 518, multi-scale classification biomarker training data 520, MIL training data 522, and post-processing training data 524.
[0168] FIG. 6 illustrates a process 600 that may be performed by the imaging-based biomarker prediction system 102, the deep learning framework 300, or the deep learning framework 402, particularly in a deep learning framework having a multi-scale configuration.
[0169] As part of the training process, in block 602, tile-labeled histopathology images are received by the deep learning framework 300. Herein, the histopathology images can be of any type, but in this example are shown as digital H&E slide images. These images can be training images of a previously determined and labeled (and therefore known) cancer type (e.g., in a supervised learning configuration). In some examples, the images can be training images of multiple different cancer types. In some examples, the images can be training images that include some or all of the images of unknown or unlabeled cancer types (e.g., in an unsupervised learning configuration). In some examples, the training images include digital H&E slide images annotated with tissue classes (to train a tile-resolution FCN tissue classifier) and other digital H&E slide images annotated with each cell. In an example of training TIL biomarker classification, each lymphocyte can be annotated in the H&E slide image (e.g., to train the annotated images of a pixel-resolution FCN segmentation classifier to train a UNet model classifier). In some examples, the training images may be digital IHC stained images for training a pixel-resolution FCN segmentation classifier, particularly images where the IHC staining targets lymphocyte markers. In some examples, the training images include images combined with molecular data, clinical data, or other annotations (such as pathway activation scores).
[0170] In the illustrated example, preprocessing is performed on the training images, such as the normalization process described herein, at block 604. Other preprocessing processes described herein may also be performed at block 604.
[0171] In block 606, the tile-labeled H&E slide training images are provided to a deep learning framework and analyzed within a machine learning construct such as a CNN, more specifically, a tile-resolution CNN implemented in some examples as an FCN model, to analyze tiled images of training images for tissue classification training, pixels of training images for cell segmentation training, and in some examples, tiled images for biomarker classification training. As a result, in block 608, a trained deep learning framework multi-scale biomarker classification model is generated, which may include a cell segmentation model and a tissue classification model. When training multiple biomarker classification models, in block 608, a separate model can be generated for each of the biomarkers TIL, PD-L1, ploidy, NC ratio, and signet ring morphology.
[0172] As a prediction process, in block 610, a new unlabeled histopathology image, such as an H&E slide image, is received and provided to a multi-scale biomarker classification model, and in block 612, the status of the biomarkers in the received histopathology image is predicted as determined by one or more biomarker classification models.
[0173] For example, in block 610, a new (unlabeled or labeled) histopathology image may be received from a physical clinical record system or primary care system and applied to a trained deep learning framework that applies its trained cell segmentation, tissue classification model, and biomarker classification model, and in block 612, a biomarker prediction score is determined. The prediction score may be determined for the entire histopathology image or for various regions of the entire image. For example, for each image, block 612 may generate the absolute number of biomarkers on the image, the percentage of cells in the tumor region associated with each of the biomarkers, and / or a designation of the biomarker classification or other information. In some examples, the deep learning framework may identify predicted biomarkers for all tissue classes identified in the image. As such, biomarker prediction probability scores may vary across images. For example, when predicting the presence of TILs, process 612 may predict the presence of TILs at different locations within the histopathology image. As a result, the prediction of TILs varies across images. This is provided as an example, and block 612 may determine any number of metrics discussed herein.
[0174] As shown in process 900 of Figure 9, in block 902, after the prediction is made, the predicted biomarker classification may be received at a report generator. In block 904, a histopathology image and therefore a clinical report for the patient may be generated, including the predicted biomarker status. In block 906, an overlay map showing the predicted biomarker status may be generated for display to a clinician or provided to a pathologist for determining a preferred immunotherapy corresponding to the predicted biomarkers.
[0175] 7 provides an exemplary process 700 for determining predicted biomarker status, and in particular for predicting TIL status. Process 700 can nevertheless be used to predict the status of any number of biomarkers and other metrics, according to the examples described herein.
[0176] The pre-processing controller receives the histopathology image and performs initial image processing (702), as described herein. In one example, the deep learning pre-processing controller receives the entire image file in any pyramidal TIFF format and identifies the edges and contours of viable tissue within the image (e.g., performs segmentation). The output of block 702 may be, for example, a binary mask of the input image, with each pixel having a value of 0 or 1, where 0 indicates background and 1 indicates foreground / tissue. The dimensions of the mask may be the dimensions of the input slide when downsampled by 128 times. This binary mask may be temporarily buffered and provided to a tiling process 704.
[0177] In process 704, the preprocessing controller applies a tissue mask process using a tiling procedure to divide the image into sub-images (i.e., tiles) that are examined separately. Because the deep learning framework is configured to run two different learning models (one for tissue classification and one for cell / lymphocyte segmentation), a different tiling procedure for each model can be run in step 704. Process 704 can generate two outputs, each including, for example, a list of coordinates defined from the top left corner of the tile. The output lists can be temporarily buffered and passed to the tissue classification and cell segmentation processes.
[0178] In the example of FIG. 7 , tissue classification is performed in process 706, which receives the pathological tissue image from process 704 and performs tissue classification on each received tile using a trained tissue classification model. The trained tissue classification model is configured to classify each tile into a different tissue class (e.g., tumor, stroma, normal epithelium, etc.). To reduce computational redundancy, process 706 can use multiple layers of tiling. For each tile, the trained tissue classification model calculates the class probability for each class stored in the model. Process 706 then determines the most likely class and assigns that class to the tile. Process 706 can output a single list of lists as a result. Each inner nested list serves as a nested classification, describing a single tile and including the tile's location, the probability that the tile is a class included in the model, and the ID of the most likely class. This information is listed for each tile. The single list of lists can be saved in an output json file of the deep learning framework pipeline.
[0179] Processes 708 and 710 perform cell segmentation and lymphocyte segmentation, respectively. Processes 708 and 710 receive the histopathology images and tile list from processes 704 and 706. Process 708 applies a trained cell segmentation model. Process 710 applies a trained lymphocyte segmentation model. That is, in the illustrated example, two pixel resolution models run in parallel for each tile in the cell segmentation tile list. In one example, both models use the UNet architecture but are trained on different training data. The cell segmentation model identifies cells and draws a border around all cells in the received tile. The lymphocyte segmentation model identifies lymphocytes and draws a border around all lymphocytes in the tile. Because hematoxylin binds to DNA, performing "cell segmentation" using digital H&E slide images is sometimes referred to as nucleus segmentation. That is, the Cell Segmentation Model process 708 performs nuclei segmentation on all cells, and the Lymphocyte Segmentation Model process 710 performs nuclei segmentation on lymphocytes.
[0180] In this example, because the same UNet architecture is used for both, processes 708 and 710 each generate two identically formatted mask array outputs. Each output is a mask array of the same shape and size as the received tile. Each array element is either 0, 1, or 2, where 0 indicates a pixel / location predicted as background (i.e., outside the object), 1 indicates a pixel / location predicted as the object boundary, and 2 indicates a pixel / location predicted as inside the object. For the output of the cell segmentation model, the object refers to a cell. For the lymphocyte segmentation model, the object refers to a lymphocyte. These mask array outputs may be temporarily buffered and provided to processes 712 and 714, respectively.
[0181] Processes 712 and 714 receive the output mask arrays of the cell segmentation (UNet) model and the lymphocyte segmentation (UNet) model, respectively. Processes 712 and 714 are performed for each received tile and are used to express the information in the mask arrays in coordinates in the coordinate space of the original whole slide image.
[0182] In one example, process 712 can access a stored image processing library and use that library to find contours around the cell interior class, i.e., contours corresponding to locations with a value of 2 in each mask. In this way, process 712 can perform a cell registration process. By matching the cell boundary class (indicated by locations with a value of 1 in each mask), separation between adjacent cell interiors is ensured. This generates a list of all contours for each mask. Next, by treating each contour as a filled polygon, process 712 determines the coordinates of the contour's centroid (center of mass), from which process 712 generates a centroid list. Next, each coordinate in the contour list and centroid list is shifted to generate output that is in the coordinate space defined by the entire received image, rather than the coordinate space specific to a single tile in the image. Without this shift, each coordinate would be in the coordinate space of the image tile that contains it. The value of each shift is the same as the coordinate of the upper left corner of the parent tile of the received image. In this example, process 714 performs the same process as process 712 for the lymphocyte class.
[0183] Processes 712 and 714 each generate a contour list output and a centroid list output corresponding to the respective UNet segmentation model. A contour is a set of coordinates that, when connected, outlines the detected object. Each contour can be represented as a line of text consisting of the constituent coordinates printed in order as pairs of numbers. Each contour list from processes 712 and 714 can be saved as a text file consisting of many such lines. A centroid list is a list of number pairs. Each of these outputs can be temporarily buffered and provided to process 716.
[0184] Process 716 receives the tissue classification output (a single list of multiple lists) from process 706, the cell centroid and contour list from process 712, and the lymphocyte centroid and contour list from process 714, and performs cell segmentation integration.
[0185] For example, process 716 can combine the paired outputs of processes 712 and 714 to produce a single concise list containing the most important information about the cells. In one example, there are two main components of process 716:
[0186] The first component of process 716 combines the information found in the cell segmentation model and the lymphocyte segmentation model. Before the information is combined, it exists as a list of cell outlines and a list of lymphocyte outlines, but the lymphocyte outlines are not necessarily a subset of cell outlines because they are the output of two independent models (712 and 714). Therefore, lymphocytes are desirably a subset of cells because (1) lymphocytes are a type of cell, and therefore if an object is not a cell, it cannot biologically be a lymphocyte, and (2) it is desired to report percentages of the dataset that have the same denominator. Therefore, the cell segmentation integration process 716 can be performed by comparing the location of each cell with the locations of all lymphocytes (in one example, this can be performed only for objects within a single tile, so that an excessive number of comparisons are not made). In one example, a cell is considered a lymphocyte only if it is "close enough" to a lymphocyte. The definition of "close enough" can be established by empirically determining the median radius of objects detected by the lymphocyte segmentation model across a set of training histopathology images. Note that this updated set of training images (e.g., 403 in Figure 4) differs from the training set of images used to train the model because this set of training histopathology images is annotated by the model itself, resulting in orders of magnitude more images—e.g., millions of automatically annotated images—that form the new or updated training set. In fact, the model's training set may continue to grow with new, incoming medical images that meet the acceptance / rejection criteria. This may be the case for tissue classification models, as well as cell segmentation and lymphocyte segmentation models. Generating a new training set from the model and evaluating subsequent images allows the model to (1) use the median value of millions of cells, rather than just those outlined in the training tiles, and (2) compare the actual size of detected objects, rather than the size of human annotations.Because lymphocyte nuclei are typically spherical, in one example, these objects were all modeled as circles (since they are two-dimensional slices of a sphere). The radii of these circles were calculated and the median was used to determine the typical size of lymphocyte detections. As a result, the final cell list will be exactly the same as the objects detected by the cell segmentation model, but the goal of the lymphocyte segmentation model is to provide a Boolean true / false label for each cell in that list.
[0187] In the second component of process 716, each cell is binned into one of the tissue classification tiles (from process 706) based on its location. Note that in the example described here, the size of the cell segmentation tile may differ from the tissue classification tile due to different model architectures. Nevertheless, process 716 is configured to determine the parent tile of each cell based on the location of the centroid, which has the coordinates of each cell centroid, the coordinates of the upper left corner of each tissue classification tile, and the size of each tissue classification tile.
[0188] Process 716 generates an output, a list of lists. Each nested inner list serves as a nested classification describing a single cell, including the coordinates of the cell's centroid, the tile number of its parent tile, the tissue class of its parent tile, and whether the cell is classified as a lymphocyte. This information is listed for each cell, and the output list is saved to the deep learning framework pipeline's output json file.
[0189] In process 718, which may be implemented by a post-processing controller such as biomarker metric processing module 326, in this particular example, any of a number of different biomarker metrics of the predicted TIL status and other TIL metrics are determined as described.
[0190] For example, process 718 can be configured to perform a tissue area calculation based on the tissue mask used in process 704 to determine the area covered by tissue. In some examples, the tissue mask is a Boolean array that takes a value of 1 where tissue is present and 0 elsewhere, so process 718 can count the number of 1s to provide a measurement of the tissue area. This value is the number of square pixels at a 128x downsampling. Multiplying this by 16,384 (i.e., for a 128*128 tissue mask) gives the number of square pixels at the native resolution (referred to as "x"). The native resolution of the image indicates the number of pixels per micron, and squaring this number gives the number of square pixels per square micron (referred to as "y"). Dividing the number of square pixels at the native resolution by this resolution scaling factor (or x / y, using the variables defined above) gives the number of square microns covered by tissue, and thus the tissue area is calculated. That is, process 718 can generate a floating point number in [0,∞] representing the tissue area in square microns, which can be used in the accept / reject model process described below.
[0191] As an example of other biomarker statistics, process 718 can be further configured to perform a total nuclei calculation using the cell segmentation integrated output from process 716. For example, the total number of nuclei on a slide is determined as the number of entries in the cell segmentation integrated output. Process 718 can also perform a % tumor nuclei calculation based on the output from this process 716. The total number of tumor nuclei on an image is the number of entries in the cell segmentation integrated output that meet the following requirements: (i) the tissue class of the parent tile is tumor, and (ii) the cells are not classified as lymphocytes.
[0192] In addition to determining biomarker statistics, process 718 can be further configured to perform an acceptance / rejection process based on the determined tissue area, total nuclei count, and tumor nuclei count. In one example, process 718 can be configured with a logistic regression model, where these three variables are used as inputs and the model's output is a binary recommendation as to whether the slide should be accepted or rejected for molecular sequencing. The logistic regression model can be trained on a training set of images using these derived variables. For example, the training images may be formed from histopathology images previously submitted for sequencing and accepted, as well as histopathology images rejected during routine pathology review. Alternatively, a preset threshold may exist, such as requiring 20% of the nuclei on a slide to be tumor, or a minimum number of tumor cells. In some examples, the model can take into account the DNA ploidy of tumor cells (data from karyotyping or DNA sequence information) and calculate an adjusted estimate of available genetic material by dividing the number of tumor nuclei by the average copy number of chromosomes detected in each tumor nucleus, by the normally expected copy number of 2. In some examples, the logistic regression model can be configured to have three possible outputs (instead of two) by adding an uncertainty zone between accept and reject that recommends manual review. For example, the penultimate output of a logistic regression model is a real number, and the final step of the model thresholds this number to 0 to generate a binary classification. Alternatively, in some examples, the uncertainty zone is defined as a range of numbers that includes 0, where values higher than this range correspond to "reject," values within the range correspond to manual review, and values lower than this range correspond to "accept." Process 718 can be configured to calculate the size of this uncertainty zone by performing cross-validation experiments. For example, process 718 can perform a training process that is repeated many times, but where each iteration uses a different random subset of images in the training set.This produces many similar, but not identical, final models, and process 718 can use this variation to determine the uncertainty range of the final logistic regression model. Thus, process 718 can produce a binary accept / reject output in some cases, and an accept / reject / manual review output in some cases.
[0193] Decisions are made using the recommendations from process 718. For example, the deep learning output post-processing controller can generate a report of images marked "accept" and automatically send those images to the genome sequencing system (112) for molecular sequencing, while images recommended as "reject" are rejected and not sent for molecular sequencing. If a "manual review" option is configured and recommended, the images can be sent to a pathologist or team of pathologists (118) to review the slide and determine whether to send for molecular sequencing or reject it.
[0194] FIG. 8 shows an exemplary process 800 that may be performed by the imaging-based biomarker prediction system 102, the deep learning framework 300, or the deep learning framework 402, particularly in a deep learning framework having a single-scale configuration.
[0195] In process 802, molecular training data is received by an imaging-based biomarker prediction system. This molecular training data is for multiple patients and can be obtained from a gene expression dataset, such as from sources described herein. In some examples, the molecular training data includes RNA-seq data. In block 804, the molecular training data is labeled with biomarkers. One form of biomarker clustering includes labeling, which can be performed by obtaining an existing label associated with a specimen, such as a tumor subtype, and associating that label with the molecular training data. Alternatively, or in addition, labeling can be performed by clustering, such as using an automated clustering algorithm. In the case of CMS subtype biomarkers, one exemplary algorithm is an algorithm for identifying CMS subtypes in the molecular training data and cluster training data according to the CMS subtype. This automated clustering can be performed, for example, within a single-class classifier module of a deep learning framework or within a slide-level labeling pipeline, such as that within deep learning framework 300. In some examples, the molecular training data received in block 802 is RNA-seq data, for example, generated using an RNA wet lab and processed using a bioinformatics pipeline.
[0196] In various embodiments, for example, each transcriptome data set can be generated by processing patient or tumor organoid samples through RNA whole exome next-generation sequencing (NGS) to generate RNA sequencing data, and the RNA sequence data can be processed through bioinformatics pipeline to generate the RNA-seq expression profile of each sample.Patient samples can be tissue samples or blood samples containing cancer cells.
[0197] RNA can be isolated from blood samples or tissue sections using commercially available reagents, such as proteinase K, TURBO DNase-I, and / or RNA Clean XP beads. The isolated RNA can be subjected to quality control protocols to determine the concentration and / or quantity of RNA molecules, including the use of fluorescent dyes and fluorescence microplate readers, standard spectrofluorometers, or filter fluorometers.
[0198] A cDNA library can be prepared from the isolated RNA, purified, and selected for cDNA molecule size selection using commercially available reagents, such as Roche KAPA HyperBeads. cDNA library preparation can include reverse transcription. In another example, a New England Biolabs (NEB) kit can be used. cDNA library preparation can include ligating adapters to the cDNA molecules. For example, UDI adapters, including Roche SeqCap dual-end adapters, or UMI adapters (e.g., full-length or partial (stubby) Y-shaped adapters) can be ligated to the cDNA molecules. In this example, the adapters are nucleic acid molecules that can serve as barcodes to identify cDNA molecules according to the sample from which they were derived and / or to facilitate downstream bioinformatics processing and / or next-generation sequencing reactions. The sequence of nucleotides within the adapters can be unique to the sample to distinguish sequencing data obtained for different samples. The adaptors may facilitate the binding of the cDNA molecules to anchor oligonucleotide molecules on the sequencer flow cell and may act as seeds for the sequencing process by providing a starting point for the sequencing reaction.
[0199] The cDNA library can be amplified and purified using reagents such as Axygen MAG PCR cleanup beads. Amplification can involve polymerase chain reaction (PCR) techniques, such as quantitative or reverse transcription-quantitative PCR (qPCR or RT-qPCR). The concentration and / or quantity of cDNA molecules can then be quantified using fluorescent dyes and fluorescence microplate readers, standard spectrofluorometers, or filter fluorometers.
[0200] Before drying under vacuum, the cDNA libraries can be pooled and treated with reagents to reduce off-target capture, such as Human COT-1 and / or IDT xGen Universal Blockers. The pools can then be resuspended in a hybridization mix, such as IDT xGen Lockdown. Probes can be added to each pool, e.g., IDT xGen Exome Research Panel v1.0 probes, IDT xGen Exome Research Panel v2.0 probes, other IDT probe panels, Roche probe panels, or other probes. Depending on the pool, the probes can be hybridized by incubating in an incubator, PCR machine, water bath, or other temperature-controlled device. The pools can then be mixed with streptavidin-coated beads or another means for capturing hybridized cDNA probe molecules, particularly cDNA molecules representing exons of the human genome. In another embodiment, polyA capture can be used. The pool can be amplified and purified again using commercially available reagents, such as the KAPA HiFi Library Amplification kit and Axygen MAG PCR cleanup beads, respectively.
[0201] cDNA libraries can be analyzed to determine the concentration or quantity of cDNA molecules, for example, by using fluorescent dyes (e.g., PicoGreen pool quantification) and a fluorescence microplate reader, standard spectrofluorometer, or filter fluorometer. cDNA libraries can also be analyzed to determine the fragment size of cDNA molecules. This can be done via gel electrophoresis techniques and may involve using a device such as the LabChip GX Touch. Pools can be cluster amplified using a kit (e.g., Illumina Paired-End Cluster Kit with PhiX Spike). In one example, the cDNA library preparation and / or whole-exome capture steps can be performed in an automated system using a liquid-handling robot (e.g., SciClone NGSx).
[0202] Library amplification can be performed on a device such as the Illumina C-Bot2, and the resulting flow cells containing the amplified target capture cDNA libraries can be sequenced on a next-generation sequencer, such as the Illumina HiSeq4000 or Illumina NovaSeq6000, to a user-selected specific on-target depth, e.g., 300x, 400x, 500x, or 10,000x. The next-generation sequencer can generate FASTQ, BCL, or other files for each patient sample or each flow cell.
[0203] When two or more patient samples are processed simultaneously on the same sequencer flow cell, reads from multiple patient samples are initially contained in the same BCL file and then split into individual FASTQ files for each patient. If the adapter sequences used for each patient sample are different, they can act as barcodes that facilitate associating each read with the correct patient sample and placing it in the correct FASTQ file.
[0204] Each FASTQ file may contain paired-end or single-read reads, and may contain short or long reads. Each read represents a single detected nucleotide sequence within an mRNA molecule isolated from a patient sample, which is inferred by using a sequencer to detect the sequence of nucleotides contained in a cDNA molecule generated from the isolated mRNA molecule during library preparation. Each read in a FASTQ file is also associated with a quality assessment, which may reflect the likelihood that an error during the sequencing procedure affected the associated read.
[0205] Each FASTQ file can be processed by a bioinformatics pipeline. In various embodiments, the bioinformatics pipeline can filter the FASTQ data. Filtering the FASTQ data can include correcting sequencer errors and removing (trimming) low-quality sequences or bases, adapter sequences, contamination, chimeric reads, overrepresented sequences, biases caused by library preparation, amplification, or capture, and other errors. Whole reads, individual nucleotides, or multiple nucleotides that may contain errors can be discarded based on a quality assessment associated with the reads in the FASTQ file, the known error rate of the sequencer, and / or a comparison of each nucleotide in the read with one or more nucleotides in other reads aligned to the same position in the reference genome. Filtering can be performed partially or entirely by various software tools. FASTQ files can be analyzed by sequencing data QC software, such as AfterQC, Kraken, RNA-SeQC, FastQC (Illumina BaseSpace Labs or see https: / / www.illumina.com / products / by-type / informatics-products / basespace-sequence-hub / apps / fastqc.html), or other similar software programs for quality control and rapid assessment of reads. In the case of paired-end reads, the reads can be merged.
[0206] For each FASTQ file, each read in the file can be aligned to the location in the reference genome whose sequence most closely matches the sequence of nucleotides in the read. Many software programs exist that are designed to align reads, including those using Bowtie, Burrows Wheeler Aligner (BWA), and the Smith-Waterman algorithm. Alignment can be directed using a reference genome by comparing the nucleotide sequence of each read to a portion of the nucleotide sequence in a reference genome (e.g., GRCh38, hg38, GRCh37, or other reference genomes developed by genome reference consortia) to determine the portion of the reference genome sequence that most likely corresponds to the sequence of the read. The alignment may also take into account RNA splice sites. Alignment can generate a SAM file, which stores the start and end positions of each read in the reference genome as well as the coverage (number of reads) of each nucleotide in the reference genome. SAM files can be converted to BAM files, BAM files can be sorted, and duplicate reads can be marked for deletion.
[0207] In one example, kallisto software can be used for alignment and quantification of RNA reads (see Nicolas L Bray, Harold Pimentel, Pall Melsted and Lior Pachter, "Near-optimal probabilistic RNA-seq quantification," Nature Biotechnology 34, 525-527 (2016), doi:10.1038 / nbt.3519). In another embodiment, quantification of RNA reads can be performed using other software, such as Sailfish or Salmon (see Rob Patro, Stephen M. Mount, and Carl Kingsford (2014) "Sailfish enables alignment-free isoform quantification from RNA-seq reads using lightweight algorithms," Nature Biotechnology (doi:10.1038 / nbt.2862) or Patro, R., Duggal, G., Love, M. Irizarry, R. A., & Kingsford, C. (2017) "Salmon provides fast and bias-aware quantification of transcript expression," Nature Methods). These RNA-seq quantification methods may not require alignment. Many software packages exist that can be used for normalization, quantitative analysis, and differential expression analysis of RNA-seq data.
[0208] For each gene, the raw RNA read count for that specific gene can be calculated. The raw read counts can be stored in a tabular file for each sample, where columns represent genes and each entry represents the raw RNA read count for that gene. In one example, the kallisto alignment software calculates the raw RNA read count for each read as the sum of the probabilities that the read aligns to the gene. Therefore, in this example, the raw counts are not integers.
[0209] The raw RNA read counts can then be normalized to correct for GC content and gene length, for example, using full quantile normalization, and the size factor method can be used to adjust sequencing depth. In one example, RNA read count normalization is performed according to the methods disclosed in U.S. Patent Application No. 16 / 581,706, entitled "Methods of Normalizing and Correcting RNA Expression Data," filed September 24, 2019, or PCT Application No. PCT19 / 52801, both of which are incorporated herein by reference in their entireties. The rationale for normalization is that the copy number of each cDNA molecule in the sequencer may not reflect the distribution of mRNA molecules in a patient sample. For example, during the library preparation, amplification, and capture steps, certain portions of mRNA molecules may be over- or under-represented due to artifacts occurring in various aspects of reverse transcription priming caused by random hexamers, amplification (PCR enrichment), rRNA depletion, and probe binding and errors generated during sequencing, which may be due to the GC content, read length, gene length, and other characteristics of each nucleic acid molecule's sequence. Each raw RNA read count for each gene can be adjusted to eliminate or reduce over- or under-representation caused by biases or artifacts in the NGS sequencing protocol. Normalized RNA read counts can be saved in a tabular file for each sample, with columns representing genes and each entry representing the normalized RNA read count for that gene.
[0210] The set of transcriptome values may refer to either normalized RNA read counts or raw RNA read counts, as described above.
[0211] 8, in block 804, the molecular training data (e.g., such RNA-seq data) is labeled with biomarkers and clustered using an automated clustering algorithm, such as an algorithm that identifies CMS subtypes in the molecular training data and an algorithm that identifies cluster training data according to CMS subtype. This automated clustering can be performed, for example, within a single-class classifier module of a deep learning framework or within a slide-level labeling pipeline such as that in deep learning framework 300.
[0212] In block 806, for each biomarker cluster (each corresponding to a different biomarker, such as a different CMS subtype or HRD), histopathology images from relevant patients are acquired. These histopathology images may be, for example, H&E slide images with slide-level labels. In block 808, for each biomarker cluster, these labeled histopathology images are provided to a deep learning framework for training a biomarker classification model, such as a multiple CMS classification model for predicting different CMS subtypes. As a result, in block 810, a set of trained biomarker classifiers (classification models) is generated. Thus, blocks 802-810 represent the training process.
[0213] The prediction process begins at block 812, where a new (unlabeled or labeled) histopathology image, such as an H&E slide image, is received and provided to the single-scale biomarker classifier generated by block 810, and at block 814, a biomarker classification on the received histopathology image is predicted as determined by one or more biomarker classification models, such as one or more CMS subtypes or HRD.
[0214] Similar to block 610, in block 814, new histopathology images are received from a physical clinical record system or primary care system, and the tissue classification model and / or biomarker classification model can be applied to the trained deep learning framework to determine biomarker predictions, which can be determined, for example, for the entire histopathology image.
[0215] 9, similar to process 600, after prediction, the predicted biomarker classification from block 814 may be received in block 902. In block 904, a histopathology image, and therefore a clinical report for the patient, may be generated, including the predicted biomarker status. In block 906, an overlay map showing the predicted biomarker status may be generated and provided to a pathologist for display to a clinician or for determining a preferred immunotherapy corresponding to the predicted biomarkers.
[0216] 10A and 10B show examples of digital overlay maps generated, for example, by the overlay map generator 324 of the system 300. These overlay maps can be generated as static digital reports displayed to a clinician or as dynamic reports that allow user interaction via a graphical user interface (GUI). FIG. 10A shows a tissue class overlay map generated by the overlay map generator 324. FIG. 10B shows a cell periphery overlay map generated by the overlay map generator 324.
[0217] In one example, the overlay map generator 324 can display the digital overlay as a transparent or opaque layer over the histopathology image, aligned so that the image location shown in the overlay and the histopathology image are in the same location on the display. The transparency of the overlay map can vary. The transparency can be user-adjustable in the overlay map generator 324's dynamic reporting mode. The overlay map generator 326 can report the percentage of labeled tiles associated with each tissue class label, the proportion of tiles classified into each tissue class, the total area of all grid tiles classified into a single tissue class, and the proportion of tiles' area classified into each tissue class. The overlay map can be displayed as a heat map showing various tissue classifications with various pixel intensity levels corresponding to various biomarker status levels. For example, in the TIL example, high intensity pixels are shown in tissue regions with a high predicted TIL status (high percentage) and low intensity pixels are shown in tissue regions with a low predicted TIL status (low percentage).
[0218] In one example, the deep learning output post-processing controller 308 can also report the total number of cells or the percentage of cells in an area defined by the user, either the entire slide, a single grid tile, all grid tiles classified as each tissue class, or cells classified as immune cells. The controller 308 can also report the number of cells classified as lymphocyte cells located within an area classified as tumor or any other tissue class.
[0219] In one example, the digital overlays and reports generated by the controller 308 can be used to help medical professionals more accurately estimate tumor purity and locate or diagnose areas of interest, including invasive tumors with tumor cells protruding into areas of non-tumor tissue surrounding the tumor. They can also assist medical professionals in prescribing treatments. For example, the number of lymphocytes in an area classified as tumor can predict whether immunotherapy will be successful in treating a patient's cancer.
[0220] In one example, the digital overlay and report generated by controller 308 can be used to determine whether a slide sample has sufficient high-quality tissue for successful genetic sequencing of the tissue, e.g., whether to implement an accept / reject / manual decision as discussed in process 700. Genetic sequencing of the tissue on a slide can be successful if the slide contains a certain amount of tissue and / or has a tumor purity value that exceeds user-defined tissue amount and tumor purity thresholds. Using process 700, controller 308 can label slides as accepted or rejected for sequencing analysis depending on the amount of tissue present on the slide and the tumor purity of the tissue on the slide. Controller 308 can also label slides as uncertain, following process 700, using a user-defined tissue amount threshold and a user-defined uncertainty range obtained from a user interacting with the digital overlay and report from generator 324.
[0221] In one example, the controller 308 implementing the process 700, for example, using the biomarker metric processing module 326, calculates the amount of tissue on a slide by measuring the total area covered by tissue in the histopathology image or by counting the number of cells on the slide. The number of cells on a slide can be determined by the number of cell nuclei visible on the slide. In one example, the controller 308 calculates the percentage of tissue that is cancerous by dividing the number of cell nuclei within a grid area labeled tumor by the total number of cell nuclei on the slide. The controller 308 can exclude cell nuclei or outer edges of cells that are located in the tumor area but belong to cells characterized as lymphocytes. The percentage of tissue that is cancerous is known as the tumor purity of the sample. The controller 308 then compares the tumor purity to a user-selected minimum tumor purity threshold and the number of cells in the digital image to a user-selected minimum cell threshold (entered by the user interacting with the overlay map generator 324). If both thresholds are exceeded, the controller approves the tissue slide depicted in the image for molecular testing, including gene sequence analysis. In one example, the user-selected minimum tumor purity threshold is 0.20, or 20%, although any number of tumor purity thresholds may be selected, including 1%, 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, or more.
[0222] In another example, the controller 308 multiplies the total area covered by tissue detected on the slide by a first multiplier value, multiplies the number of cells counted on the slide by a second multiplier value, and gives the image a composite tissue content score that sums the products of these multiplications.
[0223] In one example, the controller 308 can calculate whether grid areas labeled as tumor are spatially integrated or dispersed among non-tumor grid areas. If the controller 308 determines that the tumor areas are spatially integrated, the overlay map generator 324 can generate a digital overlay of a recommended cutting boundary that separates image areas classified as tumor and non-tumor, or image areas within non-tumor areas that are proximal to areas classified as tumor. The recommended cutting boundary can be a guide to assist a technician in dissecting a slide to separate the maximum amount of tumor or non-tumor tissue from the slide, particularly for molecular testing, including gene sequence analysis.
[0224] In one example, the controller 308 may include a clustering algorithm that calculates and reports information about the spacing and density of type-classified cells, tissue-classified tiles, or visually detectable features on a slide. The spacing information includes distribution patterns and heat maps of lymphocytes, immune cells, tumor cells, or other cells. These patterns may include clustered, dispersed, dense, and absent. This information can help determine whether immune cells and tumor cells cluster together and what percentage of cluster areas overlap, which can facilitate predicting immune infiltration and patient response to immunotherapy.
[0225] The controller 308 can also calculate and report the average tumor cell circularity, the average tumor cell perimeter, and the average tumor nuclei density.
[0226] The interval information also includes the level of intermixing of tumor and immune cells: clustering algorithms can calculate the probability that two adjacent cells on a given slide or within a region of a slide are either two tumor cells, two immune cells, or one tumor cell and one immune cell.
[0227] The clustering algorithm can also measure the thickness of some stromal patterns located around areas classified as tumors, and the thickness of the stroma surrounding this tumor region may be a predictor of patient response to treatment.
[0228] In one example, the controller 308 can also calculate and report statistics, including mean, standard deviation, sum, etc., for red, green, blue (RGB) values, brightness, hue, saturation, grayscale, and stain deconvolution information within each grid tile, either from a single slide image or aggregated from many slide images. Deconvolution involves the removal of visual signals created by several individual stains or combinations of stains, including hematoxylin, eosin, or IHC stains.
[0229] The controller 308 can also incorporate known mathematical formulas from the fields of physics and image analysis to calculate basic visually detectable features of each grid tile. Basic visually detectable features, including lines, alternating brightness patterns, and outlineable shapes, can be combined to create complex visually detectable features, including cell size, cell circularity, cell shape, and staining patterns, referred to as texture features.
[0230] In other examples, the digital overlays, reports, statistics, and estimates generated by the overlay map generator 324 may be useful in predicting patient survival, patient response to specific cancer treatments, PD-L1 status of tumor or immune clusters, microsatellite instability (MSI), tumor mutation burden (TMB), and tumor origin when the tumor origin is unknown or the tumor is metastatic. The biomarker metric processing module 326 may also calculate quantitative measures of predicted patient survival, patient response to specific cancer treatments, PD-L1 status of tumor or immune clusters, microsatellite instability (MSI), and tumor mutation burden (TMB).
[0231] In one example, the controller 308 can calculate the relative density of each type of immune cell across the entire slide in an area designated as tumor or another tissue class, including lymphocytes, cytotoxic T cells, B cells, NK cells, macrophages, etc.
[0232] In one example, the act of scanning or otherwise digitally capturing a histopathology slide automatically triggers the deep learning framework 300 to analyze the digital image of the histopathology slide.
[0233] In one example, the overlay map generator 324 allows a user to edit the cell edge or boundary between two tissue classes on a tissue class overlay map or a cell edge overlay map and save the modified map as a new overlay.
[0234] 11 illustrates a process 1100 for preparing digital images of pathology slides for tissue classification, biomarker detection, and mapping analysis that may be implemented using system 300. Process 1100 may be performed on each received image for analysis and biomarker prediction. In some examples, process 1100 may be performed in whole or in part on initially received training images. Each of the processes described in FIG. 9 may be performed by preprocessing controller 302, where any one or more of the processes may be performed by normalization module 310 and / or tissue detector 314.
[0235] In one example, such as when training a classifier model, each digital image file received at 1102 by the preprocessing controller 302 contains multiple versions of the same image content, each version at a different resolution. The file stores these copies in stacked layers, arranged by resolution, with the highest resolution image containing the most bytes at the bottom. This is known as a pyramidal structure. In one example, the highest resolution image is the highest resolution achievable by the scanner or camera that created the digital image file.
[0236] In one example, each digital image file also includes metadata indicating the resolution of each layer. The pre-processing controller 302 can detect the resolution of each layer in this metadata and compare it to a user-selected resolution standard to select the layer with the optimal resolution for analysis, in process 1104. In one example, the optimal resolution is 1 pixel per micron (downsampled by 4).
[0237] In one example, the preprocessing controller 302 receives a tagged image file format (TIFF) file with a lowest resolution of 4 pixels per micron. This 4 pixels per micron resolution corresponds to the resolution achieved by a microscope objective with a magnification of 40x. In one example, the area where tissue may be present on the slide is up to 100,000 x 100,000 pixels in size.
[0238] In one example, a TIFF file has about 10 layers, each with half the resolution of the layer below it. If the resolution of a high-resolution layer is 4 pixels per micron, then the resolution of the layer above it will be 2 pixels per micron. The area represented by one pixel in the upper layer is the size of the area represented by four pixels in the lower layer, meaning that the length of each side of the area represented by one upper-layer pixel is twice the length of each side of the area represented by one lower-layer pixel.
[0239] Each layer may be downsampled by twice the amount of the layer below it, as performed in process 1106. Downsampling is a method by which a new version of an original image can be created at a lower resolution value than the original image. There are many methods known in the art for downsampling, including nearest neighbor, bilinear, Hermite, Bell, Mitchell, bicubic, and Lanczos resampling.
[0240] In one example, downsampling by 2x means that the red, green, and blue (RGB) values of three of four pixels in a square in the high-resolution layer are replaced with the RGB value of the fourth pixel, creating a new, larger pixel in the layer above that occupies the same space as the four averaged pixels.
[0241] In one example, the digital image file does not contain a layer or image at the optimum resolution, in which case the pre-processing controller 302 can receive an image from the file having a higher resolution than the optimum resolution and downsample the image at a ratio that achieves the optimum resolution in process 1106.
[0242] In one example, the optimum resolution is 2 pixels per micron, or a "20x" magnification, but the bottom layer of the TIFF file is 4 pixels per micron, with each layer downsampled by 4x compared to the layer below it. In this case, the TIFF file has one layer at 40x magnification, the next layer at 10x magnification, but no layer at 20x magnification. In this example, preprocessing controller 302 reads the metadata, compares the resolution of each layer to the optimum resolution, and does not find a layer with the optimum resolution. Instead, preprocessing controller 302 finds the layer with 40x magnification and then downsamples the image in that layer by 2x downsampling to create an image with an optimum resolution of 20x magnification.
[0243] Also in process 1106, after acquiring the image at optimal resolution, the pre-processing controller 302 identifies all portions of the image that depict tumor sample tissue and digitally removes debris, pen marks, and other non-tissue objects.
[0244] In one example, also in process 1106, the preprocessing controller 302 distinguishes between tissue and non-tissue regions of the image and uses Gaussian deblurring to edit out pixels with non-tissue objects. In one example, any control tissue on the slide that is not part of the tumor sample tissue can be detected and labeled as control tissue by a tissue detector or manually labeled by a human analyst as control tissue to be excluded from downstream tile grid projections.
[0245] Non-tissue objects include artifacts, markings, and debris in the image, including keratin, severely compressed or pulverized tissue that cannot be analyzed visually, and objects that were not collected with the sample.
[0246] In one example, also in process 1106, the slide image contains marker ink or other writing that is detected and digitally removed by controller 302. The marker ink or other writing may be transparent on the tissue, meaning that the tissue on the slide may be visible through the ink. Because the ink for each marking is one color, the ink causes a consistent shift in the RGB values of pixels containing stained tissue underneath the ink compared to pixels containing stained tissue without the ink.
[0247] In one example, also in process 1106, controller 302 identifies portions of the slide image that contain ink by detecting portions that have RGB values that are different from the RGB values of the remainder of the slide image, and the difference in RGB values from the two portions is consistent. The tissue detector can then subtract the difference between the RGB values of the pixels in the ink portions and the pixels in the non-ink portions from the RGB values of the pixels in the ink portions to digitally remove the ink.
[0248] In one example, also in process 1106, the controller 302 eliminates pixels in the image that have low local variation. These pixels represent artifacts, markings, or blurred areas caused by an out-of-focus tissue slice, an air bubble trapped between two glass layers of a slide, or a pen mark on the slide.
[0249] In one example, also in process 1106, the controller 302 removes these pixels by converting the image to a grayscale image and then passing the grayscale image through a Gaussian blur filter to mathematically adjust each pixel's original grayscale value to a blurred grayscale value, creating a blurred image. Other filters can be used to blur the image. Then, for each pixel, the controller 302 subtracts the blurred grayscale value from the original grayscale value to create a difference grayscale value. In one example, if the pixel's difference grayscale value is less than a user-defined threshold, it may indicate that the blur filter did not significantly change the original grayscale value and that the pixel in the original image was in a blurred region. The difference grayscale value can be compared to the threshold to create a binary mask indicating where the blurred regions are, which can be designated as non-tissue regions. The mask can be a copy of the image, with the colors, RGB values, or other values within the pixels adjusted to indicate the presence or absence of a particular type of object and the location of all objects of that type. For example, a binary mask may be generated by setting the binary value of each pixel to 0 if the pixel has a differential grayscale value below a user-defined blur threshold, and setting the binary value of each pixel to 1 if the pixel has a differential grayscale value equal to or greater than the user-defined blur threshold. Areas of the binary mask with pixel binary values of 0 indicate blurred areas of the original image that may be designated as non-tissue.
[0250] The controller 302 can also mute or remove extreme brightness or darkness in the image in process 1108. In one example, the controller 302 converts the input image to a grayscale image, with each pixel receiving a numerical value corresponding to the pixel's brightness. In one example, the grayscale values range from 0 to 255, where 0 represents black and 255 represents white. For pixels with grayscale values above the brightness threshold, the tissue detector replaces the grayscale values of those pixels with the brightness threshold. For pixels with grayscale values below the darkness threshold, the tissue detector replaces the grayscale values of those pixels with the darkness threshold. In one example, the brightness threshold is approximately 210. In one example, the darkness threshold is approximately 45. The tissue detector stores the image with the new grayscale values in a data file.
[0251] In one example, the controller 302 analyzes the modified image for artifacts, debris, or markings remaining after the initial analysis in process 1110. A tissue detector scans the image and classifies remaining groups of pixels with a particular color, size, or smoothness as non-tissue.
[0252] In one example, a slide has H&E staining, and most of the tissue in the histopathology image has pink staining. In this example, the controller 302 classifies all objects that do not have a pink or red hue, as determined by the RGB values of the pixel representing the object, as non-tissue. The tissue detector 314 can interpret any color or lack of color in a pixel to indicate the presence or absence of tissue in that pixel.
[0253] In one example, the controller 302 detects the outline of each object in the image to measure its size and smoothness. Very dark pixels may be debris, and very bright pixels may be background, both of which are non-tissue objects. Therefore, the controller 302 can detect the outline of each object by converting the image to grayscale, comparing the grayscale value of each pixel to a user-determined range of values that are neither too light nor too dark, and determining whether the grayscale value is within the range to generate a binary image in which each pixel is assigned one of two numeric values.
[0254] For example, to threshold an image, controller 302 can compare the grayscale value of each pixel to a user-defined range of values and replace each grayscale value outside the user-defined range with a value of 0, and each grayscale value within the user-defined range with a value of 1. Controller 302 then draws all the outlines of all objects as the outer perimeter of each group of adjacent pixels with a value of 1. Closed outlines indicate the presence of an object, and controller 302 measures the area within each object's outline to measure the object's size.
[0255] In one example, tissue objects on a slide are unlikely to come into contact with the outer edge of the slide, and the controller 302 classifies all objects that come into contact with the edge of the slide as non-tissue.
[0256] In one example, after measuring the size of each object, the controller 302 ranks the sizes of all objects and designates the largest value as the size of the largest object. The controller 302 divides the size of each object by the size of the largest object and compares the resulting size quotient with a user-defined size threshold. If the size quotient of an object is less than the user-defined size threshold, the controller 302 designates the object as non-tissue. In one example, the user-defined size threshold is 0.1.
[0257] Before measuring the size of each object, in process 1106, the controller 302 can first downsample the input image to reduce the likelihood that portions of a tissue object will be designated as non-tissue. For example, a single tissue object may appear as a first tissue object portion surrounded by one or more additional tissue object portions having smaller sizes. After thresholding, the additional tissue object portions may have a size quotient smaller than a user-defined size threshold and may be erroneously designated as non-tissue. Downsampling before thresholding may result in a small group of adjacent 1-value pixels surrounded by 0-value pixels in the original image being included in a large group of proximal 1-value pixels. The opposite may also be true: a small group of adjacent 0-value pixels surrounded by 1-value pixels in the original image being included in a large group of proximal 0-value pixels.
[0258] In one example, the controller 302 downsamples an image with a magnification of 40x by a ratio of 16x, so that the resulting downsampled image has a magnification of 40 / 16x, and each pixel in the downsampled image represents 16 pixels of the original image.
[0259] In one example, in process 1110, controller 302 detects the boundary of each object on the slide as a cluster of pixels with a binary or RGB value not equal to zero and surrounded by pixels with an RGB value equal to zero, indicating the boundary of the object. If the pixels forming the boundary are relatively straight, controller 302 classifies the object as non-tissue. For example, controller 302 uses a closed polygon to outline a shape. If the number of vertices in the polygon is less than a user-defined minimum vertex threshold, the polygon is considered to be a simple inorganic shape that is too smooth and is marked as non-tissue. Next, controller 302 applies a tiling process to the normalized image in process 1112.
[0260] 12A-12C illustrate an exemplary architecture 1200 that may be used for the classification model of module 306. For example, the same architecture 1200 may be used for each of the tissue segmentation model 322 and the tissue classification model 320, both implemented using the FCN configurations or any neural network described herein. The tissue classifier module 306 includes a tissue classification algorithm (see FIGS. 12A-12C) that assigns a tissue class label to the image represented by each received tile (an exemplary tile 1302 is labeled in the first portion 1304 of the histopathology image 1300 shown in FIG. 13). In one example, the overlay map generator 324 may report the assigned tissue class label associated with each small square tile by displaying a grid-based digital overlay map in which each tissue class is represented by a unique color (see FIG. 12A).
[0261] A small tile size may increase the time required for the tissue classifier module 306 to analyze an input image. Alternatively, a large tile size may increase the likelihood that a tile contains more than one tissue class, making it difficult to assign a single tissue class label to the tile. In this case, instead of calculating a higher probability that one of the tissue class labels describes the image in a small square tile compared to the other tissue class labels, the architecture 1200 may calculate an equal probability that two or more tissue class labels are accurately assigned to a single small square tile.
[0262] In one example, each side of each small square tile is approximately 32 microns long, allowing approximately 5-10 cells to fit in each small square tile. This small tile size allows the tissue classifier module 306 to create more spatially accurate boundaries when determining the boundary between two adjacent small square tile regions representing two different tissue classes. In one example, each side of the small square tile can be as short as 1 micron.
[0263] In one example, the size of each tile can be set by the user to contain a specific number of pixels. In this example, the resolution of the input image will determine the length of each tile side measured in microns. At different resolutions, the tile sides may have different lengths in microns, and each tile may have a different number of cells.
[0264] The architecture 1200 recognizes various pixel data patterns in the portion of the digital image located within or near each small square tile and assigns a tissue class label to each small square tile based on those detected pixel data patterns. In one example, a medium square tile centered on a small square tile contains an area of the slide image that is sufficiently close to the small square tile to contribute to the label assignment of that small square tile.
[0265] In one example, each side of a medium square tile is approximately 466 microns long, and each medium square tile contains approximately 225 (15x15) small square tiles. In one example, this medium tile size increases the likelihood that a structural tissue feature can fit within a single medium tile, providing context to the algorithm when labeling the central small square tile. Structural tissue features may include glands, ducts, blood vessels, immune clusters, etc.
[0266] In one example, this medium tile size is chosen to counteract the shrinkage that occurs during convolution.
[0267] During convolution by architecture 1200, an input image matrix is multiplied by a filter matrix to create a result matrix. Shrinkage refers to when the result matrix is smaller than the input image matrix. The dimension of the filter matrix of a convolutional layer affects the number of rows and columns lost by shrinkage. The total number of matrix entries lost by shrinkage by processing an image through a particular CNN can be calculated depending on the number of convolutional layers in the CNN and the dimension of the filter matrix of each convolutional layer (see Figures 12A-12C).
[0268] In the example shown in Figure 12B, the combined convolutional layers lose a total of 217 matrix rows or columns from the top, bottom, and two side edges of the matrix, so the medium square tile is set to be equal to the small square tile plus 217 pixels on either side of the small square tile.
[0269] In one example, two adjacent small square tiles share a side and are each centered on a medium square tile. The two medium square tiles overlap. Of the 466*466 small pixels in each medium square tile, the two medium square tiles share all but 32*466 pixels. In one example, each convolutional layer of the algorithm (see Figures 12A and 12B) simultaneously analyzes both medium square areas, such that the algorithm generates two vectors of values (one for each of the two small square tiles).
[0270] The value vector includes a probability value for each tissue class label, indicating the likelihood that the small square tile represents that tissue class. The value vectors may be arranged in a matrix to form a three-dimensional probability data array. The location of each vector in the three-dimensional probability data array relative to other vectors corresponds to the location of the associated small square tile relative to other small square tiles included in the algorithmic analysis.
[0271] In this example, of the 466x466 (217, 156) pixels in each medium square tile, 434x434 (188, 356) is common to both medium square tiles. By analyzing both medium square tiles simultaneously, the algorithm increases efficiency.
[0272] In one example, architecture 1200 can further increase efficiency by analyzing large tiles formed by multiple overlapping medium square tiles, each containing many small square tiles surrounding one central small square tile that receives a tissue class label. In this example, the algorithm generates a data structure in the form of a three-dimensional probability data array containing one vector of probabilities for each small square tile, with the location of the vector in the three-dimensional array corresponding to the location of the small tile within the large tile.
[0273] The architecture 1200 stores this three-dimensional probability data array, for example, in the tissue classifier module 306, and the overlay map generator 324 converts the tissue class label probability for each small square tile into a tissue class overlay map. In one example, the overlay map generator 324 can compare the probabilities stored in each vector to determine the largest probability value associated with each small square tile. The tissue class label associated with that largest value can be assigned to that small square tile, and only the assigned label will be displayed in the tissue class overlay map.
[0274] In one example, the matrices generated by each layer of architecture 1200 for large square tiles are stored in graphics processing unit (GPU) memory. The capacity of the GPU memory and the amount of GPU memory required for each entry in the three-dimensional probability data array may determine the maximum possible size of the large square tiles. In one example, the GPU memory capacity is 250 MB, and each entry in the matrix requires 4 bytes of GPU memory. This allows for a large tile size of 4,530 pixels by 4,530 pixels, calculated as follows: 4 bytes / entry * 4530 * 4530 * 3 entries for each large tile = 246 (approximately 250) MB of GPU memory per large square tile. In another example, each entry in the matrix requires 8 bytes of GPU memory. In this example, a 16 GB GPU can process 32 large tiles simultaneously, with each large tile having dimensions of 4,530 pixels by 4,530 pixels, calculated as follows: 32 large tiles * 8 bytes / entry * 4530 * 4530 * 3 entries for each large tile = 14.7 (approximately 16) GB of GPU memory required.
[0275] In one example, each entry in the three-dimensional probability data array is a data entry in single-precision floating-point format (float32).
[0276] In one example, there are 16,384 (1282) non-overlapping small square tiles that form a large square tile. Each small square tile is the center of a medium square tile, with sides each approximately 466 pixels long. The small square tiles form the central region of the large square tile, with sides each approximately 4,096 pixels long. The medium square tiles all overlap, creating a border approximately 217 pixels wide around all four sides of the central region. Including the border, each large square tile has sides each approximately 4,530 pixels long.
[0277] In this example, the large square tile size allows for simultaneous calculations and reduces the percentage of redundant calculations by 99%. This can be calculated as follows: First, a pixel inside the large square tile (any pixel at least 434 pixels from the edge of the large square tile) is selected, and a region the size of a medium square tile (466 pixels per edge) is created centered around this model pixel. Next, for a small square tile at the center of this constructed region, the model pixel is contained within the corresponding medium square tile of that small square tile. There are approximately (466 / 32)^2 = 217 small square tiles within the large square tile. For pixels not inside the large square tile, the number of small square tiles that meet this condition is smaller. As the distance between the selected small square tile and the edge of the large square tile decreases, the number decreases linearly. Then, as the distance between the selected small square tile and the corner decreases, a small number of pixels (approximately 0.005%) contribute only to the classification of a single small square tile. Performing the classification on a single large square tile means that the calculation for each pixel is performed only once, rather than once for each small square tile. Thus, redundancy is reduced by a factor of approximately 217. In one example, a slide may contain several large square tiles, each of which may slightly overlap adjacent tiles, so redundancy is not completely eliminated.
[0278] An upper bound can be set on the percentage of redundant calculations (minor deviations from this limit depend on the number of large square tiles needed to cover the tissue and the relative placement of these tiles). The percentage of redundancy is 1-1 / r, where r is the redundancy ratio, and r can be calculated as (T / N+1)(sqrt(N)*E+434)^2 / (sqrt(T)*E+434)^2, where T is the total number of small square tiles on the slide, N is the number of small square tiles per large square tile, and E is the border size of the small square tiles.
[0279] FIG. 12A illustrates an example of the layer structure of architecture 1200. FIG. 12B illustrates an example of the different layers and resulting sublayer output sizes of architecture 1200, showing a tile-resolution FCN configuration. As shown, the tile-resolution FCN configuration included in tissue classifier module 306 has an additional layer of 1×1 convolution with skip connections, 8x downsampling with skip connections, a belief map layer, replacing the average pooling layer with a concatenation layer, and replacing the fully connected FCN layer with a 1×1 convolution and softmax layer. The additional layers transform the classification task into a classification-segmentation task. This means that instead of receiving and classifying the entire image as a single tissue class label, the additional layers enable the tile-resolution FCN to classify each small tile in a user-defined grid as a tissue class.
[0280] These addition and substitution layers convert the CNN into a tile-resolution FCN without the need for upsampling, which is performed in later layers of a traditional pixel-resolution FCN. Upsampling is a method that can create a new version of the original image with higher resolution values than the original. However, upsampling is a time-consuming and computationally intensive process that can be avoided in our architecture.
[0281] There are many methods known in the art for upsampling, including nearest neighbor, bilinear, Hermite, Bell, Mitchell, bicubic, and Lanczos resampling. In one example, 2x upsampling means that a pixel with red, green, and blue (RGB) values is divided into four pixels, and the RGB values of the three new pixels may be selected to match the RGB values of the original pixel. In another example, the RGB values of the three new pixels may be selected as the average of the RGB values from the original pixel and the pixels adjacent to the adjacent pixels.
[0282] Upsampling can introduce errors into the final image overlay map generated by overlay map generator 224 because the RGB values of the new pixels may not accurately reflect the visible tissue of the original slide captured by the digital slide image.
[0283] In one example, instead of labeling individual pixels, a tile-resolution FCN is programmed to analyze a large square tile made up of smaller square tiles, generating a 3D array of values, each representing the probability that a tissue class classification label matches the tissue class depicted in each smaller tile. A convolutional layer, as known in the art, performs multiplication of at least one input image matrix by at least one filter matrix. After the first convolution, the input image matrix contains values for every pixel in the large square tile input image, representing the visual data for that pixel (e.g., a value between 0 and 255 for each RGB channel).
[0284] The filter matrix may have user-selected dimensions and may contain weight values selected by the user or determined by backpropagation during training of the CNN model. In one example, in the first convolutional layer, the filter matrix has dimensions of 7x7 and has 64 filters. The filter matrix may represent visual patterns that can distinguish one tissue class from another.
[0285] In one example where RGB values are input to the input image matrix, the input image matrix and filter matrix are three-dimensional (see FIG. 12C). Each filter matrix is multiplied by each input image matrix to generate a result matrix. All result matrices generated by the filters of one convolutional layer can be stacked to create a three-dimensional result matrix with dimensions such as row, column, and depth. The final dimension of the 3D result matrix, the depth, has a depth equal to the number of filter matrices. The result matrix from one convolutional layer becomes the input image matrix of the next convolutional layer.
[0286] Returning to Figure 12A, convolutional layer titles that include " / n" (where n is a number) indicate that there is downsampling (known as pooling) of the result matrix produced by that layer. The n indicates the factor by which the downsampling occurs. Downsampling by 2 means that a downsampled result matrix with half the rows and half the columns of the original result matrix is created by replacing the square of four values in the result matrix with one of those values or a statistic calculated from those values. For example, the minimum, maximum, or average of the values may be replaced with the original value.
[0287] Architecture 1200 also adds skip connections (shown in FIG. 12A as black lines with arrows connecting the blue convolutional layers directly to the concatenation layers). The skip connection on the left involves 8x downsampling, while the skip connection on the right involves two convolutional layers that multiply the input image matrix by a filter matrix, each with dimensions of 1x1. Because the filter matrices in these layers are 1x1, only individual small square tiles contribute to the corresponding probability vectors in the result matrix created by the purple convolutional layer. These result matrices represent a small field of view.
[0288] In all other convolutional layers, the large dimensions of the filter matrices allow pixels in each medium square tile, including the small square tile in the center of a medium square tile, to contribute to the probability vector in the result matrix corresponding to that small square tile. These result matrices allow the contextual pixel data patterns surrounding the small square tile to influence the probability that each tissue class label is applied to the small square tile. These result matrices represent a large field of view.
[0289] The skip-connection 1x1 convolutional layer allows the algorithm to consider the pixel data pattern of the small central square tile as more or less important than the remaining pixel data patterns of the surrounding medium-sized square tiles. This is reflected by the weights by which the trained model is multiplied by the final result matrix from the skip-connection layer (shown on the right side of FIG. 12A) compared to the weights by which the trained model is multiplied by the final result matrix from the medium-sized tile convolutional layer between the connection layers (shown in the middle column of FIG. 10A).
[0290] The downsampling skip connection shown on the left side of Figure 12A creates a result matrix with a depth of 64. A 3x3 convolution layer with 512 filter matrices creates a result matrix with a depth of 512. A 1x1 convolution layer with 64 filter matrices creates a result matrix with a depth of 64. All three of these result matrices have the same number of rows and columns. The concatenation layer concatenates these three result matrices to form a final result matrix with the same number of rows and columns as the three concatenated matrices, and a depth of 64 + 512 + 64 (640). This final result matrix combines the large and small focus of the view matrix.
[0291] The final result matrix can be flattened to two dimensions by multiplying every entry by a coefficient and summing the products along each depth. Each element can be selected by the user or by backpropagation during model training. Flattening does not change the number of rows and columns of the final result matrix, but it does change the depth to 1.
[0292] The 1x1 convolutional layer receives the final result matrix and filters it with one or more filter matrices. The 1x1 convolutional layer can include one filter matrix associated with each tissue class label in the trained algorithm. This convolutional layer generates a 3D result matrix with a depth equal to the number of tissue class labels. Each depth corresponds to one filter matrix, and there may be a probability vector for each small square tile along the depth of the result matrix. This 3D result matrix is a three-dimensional probability data array, and the 1x1 convolutional layer stores this 3D probability data array.
[0293] A softmax layer can create a two-dimensional probability matrix from a 3D probability data array by comparing all values of each probability vector, selecting the tissue class associated with the maximum value, and assigning that tissue class to the small square tile associated with that probability vector.
[0294] The stored three-dimensional probability data array or 2D probability matrix can then be converted into a tissue class overlay map in the final confidence map layer of Figure 10A to efficiently assign tissue class labels to each tile.
[0295] In one example, to counteract the shrinkage, the input image matrix adds rows and columns to all four outer edges of the matrix, with each value entry in the added rows and columns equal to zero. These rows and columns are called padding. In this case, the training data input matrix would have the same number of added rows and columns with value entries equal to zero. Any difference in the number of padded rows or columns in the training data input matrix would result in values in the filter matrix that would not cause the tissue class locator 216 to accurately label the input image.
[0296] In the FCN shown in Figure 12A, due to the gray and blue layers, a total of 217 outer rows or columns on each side of the input image matrix are lost due to shrinkage before skip connections. Only pixels in the small square tiles have vectors corresponding to the resulting matrix created after the green layer.
[0297] In one example, each medium square tile is not padded by adding rows and columns with zero-valued entries around the input image matrix corresponding to each medium square tile, because the zeros replace image data values from adjacent medium square tiles that the tissue class locator 216 needs to analyze. In this case, the training data input matrix is also not padded.
[0298] FIG. 12C is a visualization of the depths of an exemplary three-dimensional input image matrix being convolved with two exemplary three-dimensional filter matrices.
[0299] In an example where the input image matrix contains RGB channels for each medium square tile, the input image matrix and filter matrices are three-dimensional: in one dimension, the input image matrix and each filter matrix has three depths: one for the red channel, one for the green channel, and one for the blue channel.
[0300] The red channel (first depth) 1202 of the input image matrix is multiplied by the corresponding first depth of the first filter matrix. The green channel (second depth) 1204 is multiplied in a similar manner, and the blue channel (third depth) 1206 is multiplied in a similar manner. The red, green, and blue product matrices are then summed to create the first depth of the three-dimensional result matrix. This is repeated for each filter matrix, creating additional depths of three-dimensional result matrices corresponding to each filter.
[0301] A wide variety of training sets can be used to train the CNN or FCN included in the tissue classifier module 306.
[0302] In one example, the training set may include JPEG images of medium-sized square tiles, each with a tissue class label assigned to its central small square tile taken from at least 50 digital images of histopathology slides at a resolution of approximately 1 pixel per micron. In one example, a human analyst delineates and labels all relevant tissue classes (annotates various tissue classes) or labels each small square tile of each histopathology slide as non-tissue or a specific type of cell. Tissue classes may include tumor, stroma, normal, immune cluster, necrosis, hyperplasia / dysplasia, and red blood cells. In one example, each side of each central small square tile is approximately 32 pixels long.
[0303] In one example, the training set images are converted into an input training image matrix and processed by the tissue classifier module 306 to assign a tissue class label to each tile image of the training images. If the tissue classifier module 306 does not accurately label the validation set of training images to match the corresponding annotations added by a human analyst, the weights of each layer of the deep learning network can be automatically adjusted by stochastic gradient descent with backpropagation until the tissue classifier module 306 accurately labels the majority of the validation set of training images.
[0304] In one example, a training dataset has multiple classes, where each class represents a tissue class. The training set generates a unique model using specific hyperparameters (e.g., number of epochs, learning rate) that can recognize the content of digital slide images and classify them into various classes. The tissue classes may include tumor, stroma, immune cluster, normal epithelium, necrosis, hyperplasia / dysplasia, and red blood cells. In one example, if there are sufficient training sets for each tissue class, the model can classify an unlimited number of tissue classes.
[0305] In one example, the training set images are converted into grayscale masks for annotation, where different values (0-255) of the mask image represent different classes.
[0306] Because each histopathology image may exhibit significant variation in visual features, including tumor appearance, the training set may contain highly diverse digital slide images to better train the model for the wide variety of slides that may be analyzed. Images in the training data may be subjected to data augmentation (including rotation, scaling, color jitter, etc.) before being used to train the model.
[0307] A training set can also be cancer type-specific, where all histopathology slides from which digital images are generated in a particular training set contain tumor samples from the same cancer type. Cancer types can include breast, colorectal, lung, pancreatic, liver, stomach, skin, etc. Each training set can generate a unique model specific to the cancer type. Each cancer type can also be divided into cancer subtypes known in the art.
[0308] In one example, the training set may be derived from a pair of histopathology slides. The pair of histopathology slides includes two histopathology slides, each containing one slice of tissue, where the two slices of tissue were located substantially proximal / nearly adjacent to each other in a tumor sample. Therefore, the two slices of tissue are substantially similar. One of the slides is stained only with H&E staining, while the other slide in the pair is stained with IHC staining for a specific molecular target. Areas on the H&E-stained slide corresponding to the areas in the slide pair where the IHC staining appears are annotated by a human analyst as containing the specific molecular target, and the tissue class locator receives the annotated H&E slide as the training set. Substantially similar slides may include, for example, a slide pair that includes an H&E-stained slide and a slide containing molecular sequence data extracted from one of the adjacent slides, or other combinations, such as a slide containing IHC staining and a slide containing molecular sequence data, or both containing similar molecular sequence data.
[0309] For example, in some embodiments, two or more samples are obtained from a subject (e.g., two or more adjacent tissue slices can be taken). In some cases, the tissue slices are acquired such that a portion of a pathology slide prepared from each slice is imaged, while a portion of the pathology slide is used to acquire sequencing information.
[0310] An appropriate training dataset can be used to train an optimization model according to embodiments of the present disclosure. In some embodiments, curating a training dataset can include collecting a series of pathology reports and associated sequencing information from multiple patients. For example, a physician can perform a tumor biopsy on a patient by extracting a small amount of tumor tissue / specimen from the patient and sending the specimen to a laboratory. In the laboratory, slides can be prepared from the specimen using slide preparation techniques such as freezing the specimen and slice layers, setting the specimen in paraffin and slice layers, smearing the specimen onto a slide, or other methods known to those skilled in the art. For purposes of the following disclosure, slides and slices can be used interchangeably. A slide stores a slice of tissue from a specimen and receives a label identifying the specimen from which the slice was extracted and the sequence number of the slice from the specimen. Traditionally, pathology slides can be prepared by staining a specimen to reveal cellular features (such as cell nuclei, lymphocytes, stroma, epithelium, or whole or some other cells). The pathology slide selected for staining is traditionally the terminal slide of a specimen block. Slicing the specimen proceeds with a series of initial slides that can be prepared for staining and diagnosis purposes. The next consecutive slice in the series can be used for sequencing, and the final terminal slide can be processed for additional staining. If the terminally stained slide is too far from the sequenced slide, another slide closer to the sequenced slide can be stained, separating the sequenced slide by the stained slide. While slight variations in slice size are expected, they are expected to be minimal because tissue is sliced at thicknesses approaching 4 um for paraffin slides and 35 um for frozen slides. Laboratories generally find that distances of less than 40 um (approximately 10 slides / slice) do not result in substantial variations in tissue slices.
[0311] If slices of a specimen vary significantly from slice to slice (which is rare), the outliers may be discarded and not processed further. The pathology slides 510 may be various stained slides taken from tumor samples from patients. Some slides and sequencing data may be taken from the same specimen to ensure data robustness, while other slides and sequencing data may be taken from each unique specimen. The more tumor samples in the dataset, the higher the accuracy can be expected from the prediction of cell type RNA profiles. In some embodiments, the stained tumor slides may be reviewed by a pathologist to identify cellular characteristics (such as the amount of cells and how they differ from normal cells of that type or a similar type).
[0312] In this case, the trained tissue classification model 320 receives a digital image of H&E stained tissue and predicts tiles that may contain IHC staining or a given molecular target, and the overlay map generator 326 generates an overlay map that indicates which tiles are likely to contain the IHC target or the given molecule. In one example, the resolution of the overlay is at the level of individual cells.
[0313] Overlays generated by a model trained on one or more training sets can be reviewed by a human analyst to annotate digital slide images and add them to one of the training sets.
[0314] The pixel data patterns that the algorithm detects may represent visually detectable features, some examples of which may include color, texture, cell size, shape, and spatial organization.
[0315] For example, the color of the slide provides contextual information. Purple areas on the slide are more likely to be densely cellular and invasive tumors. Tumors also make the surrounding stroma more fibrotic in a desmoplastic reaction, making the normally pink stroma appear blue-gray. Color intensity also helps identify individual cells of a particular type (e.g., lymphocytes are uniformly very dark blue).
[0316] Texture refers to the distribution of staining within cells. Most tumor cells have a rough, heterogeneous appearance, with bright pockets in the nucleus and dark nucleoli. A zoomed-out view of many tumor cells results in this rough appearance. Many non-tumor tissue classes each have characteristic features. Furthermore, the pattern of tissue classes present in a region can indicate the type of tissue or cellular structure present in that region.
[0317] Additionally, cell size often indicates tissue class: if a cell is several times larger than normal cells elsewhere on the slide, it is more likely to be a tumor cell.
[0318] The shape of individual cells, particularly how round they are, can indicate what type of cell they are. Fibroblasts (stromal cells) are usually elongated and thin, while lymphocytes are very round. Tumor cells may be more irregularly shaped.
[0319] The organization of groups of cells can also indicate tissue class. Normal cells are often organized in structured, recognizable patterns, while tumor cells grow in denser, more disorganized clusters. Each type and subtype of cancer may produce tumors with specific growth patterns, including the location of cells relative to tissue features, the spacing of tumor cells relative to each other, the formation of geometric elements, etc.
[0320] The techniques herein can be extended to other architectures. For example, Figure 14 shows an imaging-based biomarker prediction system 1400 that also uses separate pipelines for tissue classification and cell classification. System 1400 can be used for determining various biomarkers, including PD-L1, as described in the Examples herein. Furthermore, like the other architectures herein, system 1400 can be configured to predict biomarker status and tumor status and tumor statistics based on 3D image analysis.
[0321] The system 1400 receives one or more digital images of a histopathology slide and creates a high-density grid-based digital overlay map that identifies the majority of classes of tissue visible within each grid tile of the image. The system 1400 can also generate a digital overlay drawing that identifies each cell in the histopathology image at the resolution level of individual pixels.
[0322] System 1400 includes a tissue detector 1402 that detects areas of a digital image that contain tissue and stores data including the locations of the areas detected as containing tissue (e.g., the location of the pixel using a reference location in the image, such as the 0,0 pixel location). Tissue detector 1402 forwards tissue area location data 1403 to a tissue class tile grid projector 1404 and a cell tile grid projector 1406. Tissue class tile grid projector 1404 receives the tissue area location data 1403 and performs tissue classification on the tiles for each of a number of tissue class labels. A tissue class locator 1408 receives the resulting tile classifications and determines where each tissue class is located in the digital image by calculating a percentage representing the likelihood that the tissue class label accurately describes the image in each tile. For each tile, all percentages calculated for all tissue class labels sum to 1, reflecting 100%. In one example, the tissue class locator 1408 assigns one tissue class label to each tile to determine where each tissue class is located in the digital image, and stores the calculated percentages and the assigned tissue class labels associated with each tile.
[0323] In one example, the system 1400 includes a multi-tile algorithm that simultaneously analyzes many tiles in an image, both individually and in combination with the portion of the image surrounding each tile. The multi-tile algorithm can provide multi-scale, multi-resolution analysis that captures both the content of individual tiles and the context of the portion of the image surrounding the tile. Because the portion of the image surrounding two adjacent tiles overlaps, analyzing many tiles and their surroundings simultaneously, rather than analyzing each tile individually with its surroundings, reduces computational redundancy and improves processing efficiency.
[0324] In one example, the system 1400 can store the analysis results in a three-dimensional probability data array, which includes one one-dimensional data vector for each analyzed tile. In one example, each data vector includes a list of percentages that sum to 100%, each indicating the probability that each grid tile contains one of the analyzed tissue classes. The position of each data vector in the orthogonal two-dimensional plane of the data array relative to the other vectors corresponds to the position of the tile associated with that data vector in the digital image relative to the other tiles.
[0325] The cell type tile grid projector 1406 receives the tissue area location data 1403, identifies and classifies cells within the tiles, and projects a cell type tile grid onto the area of the image containing the tissue. The cell type locator 1410 detects each biological cell of the digital image within each grid, outlines the outer edge of each cell, and can classify each cell by cell type. The cell type locator 1410 stores data including each pixel containing the location and outer edge of each cell, as well as the cell type label assigned to each cell.
[0326] The overlay map generator and metric calculator 1412 can retrieve the stored three-dimensional probability data array from the tissue class locator 1408 and convert it into an overlay map that displays the tissue class label assigned to each tile. The tissue class assigned to each tile may be displayed as a transparent color unique to each tissue class. In one example, the tissue class overlay map displays the probability of each grid tile for a user-selected tissue class. The overlay map generator and metric calculator 1412 also retrieves stored cell location and type data from the cell type locator 1410 and calculates metrics related to the number of cells in the entire image or within a tile assigned to a particular tissue class.
[0327] FIG. 15A shows an overview of an exemplary process 1500 implemented by the imaging-based biomarker prediction system 1400, illustrating the model inference pipeline for predicting biomarkers in this exemplary PD-L1 model. In process 1500, many tiles are processed in parallel, leveraging the fully convolutional model architecture of system 1400. In one example, using a GeForce GTX 1080 Ti GPU and a 6th Generation Intel® Core™ i7 processor, process 1500 required 2.8 seconds to classify a single 4096 x 4096 pixel image. Process 1500 also included tissue detection and artifact removal algorithms to function in a fully automated manner in a real-world setting where slides may contain artifacts.
[0328] In the first process 1502, initial tissue segmentation is performed by the tissue detector 1402, e.g., applying a tissue masking algorithm to automatically outline the tissue (red outline) to generate a bounding box (not shown) around the tissue of interest. Aligned with the upper left corner of the bounding box, the tissue region is divided into large, non-overlapping 4096 x 4096 input windows (blue dashed lines). Typically, 10 to 30 input windows are required to cover the tissue. Areas of the large windows that extend beyond the bounding region are filled with zeros (gray areas).
[0329] In process 1504, trained classification model prediction is performed. In the illustrated example, the large input window contained 128 × 128 = 16,384 smaller 32 × 32 tiles (the grid is much finer than shown). The large input window was zero-padded on all sides (length 217) to account for the overlapping 466 × 466 tile edges at the center of each 32 × 32 smaller tile. Each large window was passed through one or more trained models 1506 of a deep learning framework (including the tissue classification process of projector 1404 and the cell classification process of projector 1406). If the trained models 1506 are fully convolutional, in one example, each tile in the large input window is processed in parallel to generate a 128 × 128 × 3 probability cube (with three classes). Each 1 × 1 × 3 vector in this probability cube corresponds to a 32 × 32 pixel region at the center of each 466 × 466 tile in the original image. The resulting probability cubes are assembled into a probability map of the entire image.
[0330] Process 1508 displays the image associated with the tissue masking step of process 1502. The assembled probability map produced by process 1504 is passed through this tissue mask to remove the background. In this example, both the background and marker areas have been removed by the masking algorithm of process 1508.
[0331] Process 1510 displays an image showing a classification map identifying one or more classified regions, such as each distinct region for each biomarker classification. In one example, the maximum probability class (argmax) is assigned to each tile via process 1508, and a classification map of three biomarker classes (PD-L1+, PD-L1-, and others) is generated in process 1510. The classification map shows each of these biomarker classes and their identified locations relative to the original histopathology image.
[0332] Process 1512 performs statistical analysis on the biomarker classifications from the classification map and displays the resulting prediction scores for the biomarkers. In this example, the number of predicted PD-L1 positive tiles is divided by the total number of predicted tumor tiles to arrive at an exemplary model score.
[0333] The tissue class locator 1408 may have an architecture similar to that of the architecture 1200. The architecture is similar to that of an FCN tile-resolution classifier. The architecture 1550 may be formed from three main components: 1) a fully convolutional residual neural network (e.g., built on ResNet-18) backbone that processes wide field-of-view (FOV) images; 2) two branches that process small FOVs; and 3) a concatenation of small and large FOV features for multi-FOV classification. The ResNet-18 backbone contains multiple shortcut connections, shown by dotted lines, and the feature maps are also downsampled by 2x. The small FOV branch appears after the second convolution block. The feature maps of the small FOV branch are downsampled by 8x to match the dimensionality of the ResNet-18 feature maps. These feature maps are concatenated before passing through a softmax output to generate a PD-L1 biomarker prediction (confidence) map.
[0334] In one example, the model backbone includes an 18-layer version of ResNet (ResNet-18) with some modifications. The ResNet-18 backbone was converted to a fully convolutional network (FCN) by removing the global average pooling layer and eliminating zero-padding in downsampled layers. This allows for the output of 2D probability maps rather than 1D probability vectors (see Figure 15B). In the illustrated example, the tile size (466 x 466 pixels) is more than twice that of a standard ResNet, providing the model with a large FOV from which it can learn morphological features of its surroundings. While the tissue class locator 1408 in this example adds multiple fields of view (multi-FOV) capabilities to the ResNet architecture, it should be understood that the tissue class locator 1408 may be comprised of a separate network architecture adapted to incorporate the multi-FOV approach disclosed herein.
[0335] The FCN configuration of architecture 1505 offers many advantages, including overcoming the accuracy degradation challenges traditionally associated with "very deep" neural networks (including those with more than 16 convolutional layers) (see, for example, He et al., "Deep Residual Learning for Image Recognition" (2015), arXiv ID:1512.03385v1, and Simonyan et al., "Very Deep Convolution Networks for Large-Scale Image Recognition" (2014), arXiv ID:1409.1556v6). Architecture 1550 contains a stack of convolutional layers interleaved with "shortcut connections" that skip intermediate layers. These shortcut connections use past layers as reference points to guide deeper layers in learning residuals between layer outputs, rather than learning identity mappings between layers. This innovation improves convergence speed and stability during training, enabling deep networks to perform better than shallow networks.
[0336] The tissue class locator 1408 may include two additional branches with receptive fields restricted to a small FOV (32 × 32 pixels) centered on the second convolutional feature map (see Figure 15B). One branch passes a copy of the small FOV through a convolutional filter, while the other branch is a standard shortcut connection with downsampling. The features generated by these additional branches are concatenated with features from the main backbone in a softmax layer just before the model output is converted to a probability. In this way, the tissue class locator 1408 combines information from multiple FOVs, similar to the various zoom levels a pathologist might encounter when diagnosing a slide. Furthermore, this architecture ensures that the central region of each tile contributes more to the classification than the edges of the tile, resulting in a more accurate classification map across the entire histopathology image.
[0337] Figure 15B illustrates an exemplary training process 1570 for the imaging-based biomarker prediction system 1400 and the generation of an overlay map output, in which the location of the PD-L1 biomarker is predicted from analysis of IHC and H&E histopathology images. During the model training process 1570, matching areas in the IHC and H&E digital images were annotated by a medical professional. However, in some cases, staining in adjacent tissue slices may be automatically annotated. Images showing PD-L1+ and PD-L1- were annotated and used to train the system. The annotated regions of the H&E images were tiled into overlapping tiles (466 x 466 pixels) with a stride of 32 pixels to generate training data. The tissue class locator 1408 was then trained using a cross-entropy loss function. The yellow box in the model schematic indicates the central region, cropped for a small FOV. The resulting PD-L1 classification model is saved in a trained deep learning framework 1574.
[0338] Figure 15B also shows an exemplary prediction process 1572 in which each image was divided into large, non-overlapping 4096 x 4096 input windows (blue dashed lines). Each large window was passed through the trained model. Because the deep learning framework 1574 is fully convolutional, each tile within the large input window was processed in parallel, generating a 128 x 128 x 3 probability cube (the last dimension represents the three classes). The resulting probability cubes were slotted into place and assembled to generate a probability map for the entire image. The class with the highest probability was assigned to each tile, and a PD-L1 prediction report was generated.
[0339] Figures 16A-16F show input histopathology images received by imaging-based biomarker prediction system 1400, the corresponding overlay map generated by system 1400 to predict the location of the IHC PD-L1 biomarker, and the corresponding IHC-stained tissue image used as a reference to determine the accuracy of the overlay map. The IHC-stained tissue images were obtained from the test cohort but were not applied to system 1200 during model training. Figures 16A-16C show a representative example of PD-L1-positive biomarker classification. Figure 16A displays the input H&E image, Figure 16B displays the probability map overlaid on the H&E image, and Figure 16C shows PD-L1 IHC staining for reference. Figures 16D-16F show a representative example of PD-L1-negative biomarker classification. Figure 16D displays the input H&E image, Figure 16E displays the probability map overlaid on the H&E image, and Figure 16F shows PD-L1 IHC staining for reference. The color bar indicates the predicted probability of tumor PD-L1+ class.
[0340] Among the benefits offered by our deep learning framework, we can improve accuracy by breaking shift invariance. Shift invariance, or uniformity, is a property of linear filters, such as convolutions, whose response does not explicitly depend on position. That is, if we shift the signal, the output image remains the same, but with the shift applied. While shift invariance is desirable for most image classification tasks (Le Cun (1989)), in our examples, it is generally undesirable for objects near the edge of the tile to contribute equally to the classification.
[0341] Figure 17 illustrates an exemplary benefit of the multi-FOV strategy described with reference to Figures 14, 15A, and 15B. The upper, large FOV (red box) contains both PD-L1+ tumor cells (purple, top left) and stroma (pink). Only the stroma is contained in the small FOV (green box). After passing through the convolutional layers of the tissue class locator 1408, the tumor area generates a unique pattern (colored squares) that differs from the pattern generated by the stromal area (white squares). After the patterns from the large and small FOV branches are concatenated, the model may predict "Other." In the lower part, the field of view shifts, and the PD-L1+ tumor area is now contained within the small FOV. This tumor area generates the same convolutional filter pattern in both the large and small FOV branches (colored squares). Concatenating the learned features, the tissue class locator 1408 is more likely to predict PD-L1+ tumors. Therefore, without the small FOV used to train the tissue class locator 1408, the system 1400 may have predicted PD-L1+ tumors for both images. Instead, the multi-FOV strategy of the architecture of FIG. 15 allows the network to preferentially classify what is in the center of the image while taking advantage of the rich contextual information in the surrounding areas.
[0342] Still other architectures can be used with any of the classifier examples herein to predict biomarker status, tumor status, and / or metrics thereof, particularly using multi-instance learning techniques.
[0343] In the example discussed herein, a classification model architecture based on the FCN architecture described in FIG. 12A was trained on digital images of histopathology slides, which may include a matrix of annotations. Training from such digital images is performed tile-by-tile. For example, only tiles with annotations (i.e., labels) are provided to the deep learning framework as training tiles. Tiles of the digital image without annotations may be discarded. Furthermore, each column and row of the matrix corresponds to a separate grid of NxM pixels in the digital image. To properly assign annotations to a digital image having multiple tiles from a matrix with columns and rows, in some examples, it may be advantageous to take column (i) and row (j) from the matrix and assign annotations to a central region of the grid starting at pixels N(i) and M(j) and extending to the next [N-1] to [M-1] pixels. Here, tiles spanned by the central region are assigned the annotations of the i, j matrix. Thus, the matrix can accurately represent annotations for each tile within the larger digital image.
[0344] The FCN architecture can use large tiles as input, while labels are obtained from a central region that maps to annotation mask points. The FCN architecture can learn from both the central region of the large tile and the pixels surrounding it, with the central region contributing more to the prediction. Furthermore, slide metadata can be stored in a feature vector, such as a vector associated with identified slides for slide-level labeling. This includes patient features that may improve model performance. Multiple tiles from a grid of digital images and their corresponding annotation matrices are sequentially provided to the FCN architecture, which is trained to classify N x M tiles according to the annotations contained in the matrix and the slides themselves according to the annotations contained in the feature vector. The output of the FCN architecture may include a matrix of predicted classifications for each tile, which can be aggregated into a vector of predicted classifications for each slide. The matrix can be converted into a digital overlay by associating the best classification for each tile with a color that can be superimposed on the corresponding grid location in the digital image. In some examples, the matrix may be converted into multiple digital overlays, each corresponding to a different classification, with associated color intensities assigned to the overlays based on the percentage of confidence associated with each classification. For example, a tile with a 30% probability of being tumor, a 50% probability of being stromal, and a 20% probability of being normal may be assigned a single overlay of stroma, the most likely tissue within the tile, or a first overlay with a first color intensity of 30% and a second overlay with a second color intensity of 50% may be assigned to identify the type of tissue the tile may comprise.
[0345] However, even classification models based solely on architectures that may not support per-tile annotations, such as architectures similar to Resnet-34 or Inception-v3, can be trained using only digital images of pathology slides that contain only a vector of annotations, where each entry in the vector is a patient feature or metadata annotation applied to the slide. In some cases, even architectures that support per-tile annotations may not have access to per-tile annotations for certain features.
[0346] To use histopathological images in training neural networks without tile annotations or to train neural networks to identify biomarkers trained on molecular training data, in some examples, the present technology includes deep learning training architectures configured for label-free training. In examples, the architectures do not require tile-level annotations when training tissue classification models. Furthermore, the label-free training architectures are neural network agnostic in that they enable label-free training that is independent of the neural network configuration (i.e., ResNet-34, FCN, Inception-v3, UNet, etc.). The architectures can analyze a set of possible training images and predict tiles to exclude from training. Thus, in some examples, images may be discarded from training, while in other examples, tiles may be discarded but the remainder of the images may be used for training. These techniques significantly reduce the amount of training data required, significantly shortening the time required for training and, in some cases, significantly shortening the time required for training updates of the classification models herein. Furthermore, because the technology does not require labeling by a pathologist, the time required to train classification models is significantly reduced and annotation errors and variability between experts are avoided.
[0347] Alternatively, in some cases, training can be performed using weakly supervised learning, which involves only image-level labeling and does not include local labeling of tissues, cells, tumors, etc. The architecture can consist of an unlabeled training front-end with an algorithm equipped with a customized cost function that selects tiles to use as inputs for a particular label. This process can be iterative, first treating each histopathology image as a collection of tiles, with the image's single label applied to all tiles in the collection. The tiles can then be applied to an inference pipeline, such as through a network such as ResNet34, Inception-v3, or FCN. Additionally, predefined tile selection criteria, such as the probability of the neural network output, can be used to select the output image tile to be provided as input to the same neural network for the next round. This process can be repeated many times, and given enough collections and tiles as inputs to the neural network, it will learn to distinguish between tiles of different classes with increasing accuracy as more iterations are performed.
[0348] FIG. 18 illustrates an exemplary machine learning architecture 1800 in an exemplary configuration for performing label-free annotation training of a deep learning framework to implement the processes described in the examples herein. The deep learning framework 1802 may be similar to other deep learning frameworks described herein with multi-scale and single-scale classification modules and includes a pre-processing and post-processing controller 1804, and the execution process is described in the similar examples of FIGS. 1 and 3. The deep learning framework 1802 includes a cell segmentation module 1806 and a tissue classifier module 1808, each of which is configured as a tile-based neural network classifier. The deep learning framework 1802 further includes multiple different biomarker classification models 1810, 1812, 1814, and 1816, each of which can be configured with different neural network architectures, some of which may have multi-scale configurations and some of which may have single-scale configurations. These different neural network architectures can be configured for training using annotated or unannotated images. Some of these architectures can be configured for training using training images with tile annotations, while others can be configured for training using training images without tile annotations. For example, some architectures may be configured to only accept slide-level annotations of images (i.e., annotations for the entire image, not annotations that identify specific tile characteristics, cell segmentation, or tissue segmentation). Examples of neural network architecture types for modules 1810-1816 include ResNet-34, FCN, Inception-v3, and UNet.
[0349] The annotated images 1818 may be provided to the deep learning framework 1802 for training of various classification modules using the techniques described above. In some examples, entire histopathology images are provided to the framework 1802 for training. In some examples, the annotated images 1818 are passed directly to the deep learning framework 1802. In some examples, the annotated images 1818 may be annotated at a reduced granularity. Thus, in some examples, the multi-instance learning (MIL) controller 1821 may be configured to further separate the annotated images 1818 into multiple tile images, each corresponding to a different portion of the digital image, and the MIL controller 1821 applies the tile images to the deep learning framework 1802. However, in the architecture 1800, unannotated images 1820 can be used for training of classification modules by first providing the images 1820 to the MIL controller 1821 with a front-end tile selection controller 1822. In some examples, the MIL controller 1821 can be configured to separate the unannotated image 1820 into multiple tile images, each corresponding to a different portion of the digital image, and the MIL controller 1821 applies the tile images to the deep learning framework 1802. In one example, the architecture 1800 deploys weakly supervised learning to train a convolutional neural network architecture (e.g., FCN, ResNet34, Inception-v3, etc.) to classify local tissue regions using only slide-level labels. Because weakly supervised learning does not require localized annotations, labeling can be performed faster, resulting in a larger set of labeled slides. Thus, this architecture 1800 can be used to train models that complement or improve FCN classification or to further train the FCN-based model itself.
[0350] In this illustrated example, the front-end tile section controller 1822 is configured with a feedback configuration that allows the tile section process to be combined with a classification model, allowing the tile section process to be informed by a neural network architecture, such as an FCN architecture. In some examples, the tile selection process performed by the controller 1822 is a trained MIL process. For example, one of the biomarker classification models 1810-1816 can generate an output during training, which is used as an initial input to guide the tile selection controller 1822. The MIL process is typically an iterative process where initial tile selection is difficult; by informing the MIL process of the controller 1822 with guidance from, for example, the FCN architecture predictions, the MIL process of the controller 1822 starts with better examples and converges more quickly to a stable and useful FCN classifier. In yet another example, combining results from an FCN (or other neural network) architecture with the MIL process of the controller 1822 may include combining only the results of the vector outputs in the concatenation layer, such that the matrix output is obtained by voting for the best. The FCN architecture and MIL process of the controller 1822 can configure the same prediction task (i.e., look for the same biomarkers). However, in some cases, the prediction outputs of the two classification processes may differ, and in such cases, combining the results by voting for the best one can be performed. In yet another example, the outputs from the MIL framework and the FCN architecture can be combined to obtain better slide-level predictions. In this case, the MIL uses a slide-level loss function as a learning criterion, and the output from the FCN architecture is used to derive a guided truth for the MIL loss calculation.
[0351] The tile section controller 1822 can be implemented in several different ways.
[0352] An example of the controller 1822's tile selection process is described with reference to a single-class basic framework. In a single-class example, tiles in a histopathology image are classified as either belonging to a target class (Class 1) or not (Class 0). Class 0 can be considered background and not in the target class. An instance-based MIL process requires the selection of tiles to be used as examples in training. For single-class problems, the controller 1822 can be configured with a trained model that returns the following classification: if a slide has no tiles that belong to the target class, all tiles should return a low inference score of zero, and the slide is labeled as Class 0 during training; and if a slide has tiles that belong to the target class, the slide is labeled as Class 1, and tiles that belong to the target class should return a high inference score of 1, while all other tiles should return a low inference score of 0. To train a classifier model (e.g., models 1810-1816), tiles that represent slide-level classes must be identified. Since we know that Class 0 slides do not have Class 1 tiles, any tile can be used as a training example. However, the best choice is to use the tile with the lowest model performance. Since all tiles must have an inference score of 0, the tiles with the highest scores should be used to train the model. These highest-scoring tiles are called the "top k" tiles, where k is an integer (e.g., 5, 10, 15) that indicates the number of tiles from that slide that are being used for training.
[0353] For a class 1 slide, controller 1822 may identify the tiles that are most likely to be of class 1. This decision can be complicated because a class 1 slide may contain tiles of both class 0 and class 1. However, for a classifier model trained on class 0 tiles of class 0 slides, this means that tiles in class 1 slides that are similar to class 0 tiles will have a lower inference score. Similarly, this means that tiles that are dissimilar to class 0 tiles should have a higher score. Thus, the tiles that are most likely to actually be class 1 are the tiles with the highest inference scores. Therefore, the top k tiles should again be selected as examples for training the model.
[0354] This means that the top k scoring tiles for both class 0 and class 1 slides should be used as training examples, as seen in the tile selection framework 1900 in Figure 19. The rows of numbers represent the inference scores for the various tiles. First, an annotated histopathology image without tiles is provided in 1902, and model inference is performed in process 1904 for each of the tiles in the image.
[0355] Framework 1900 can be used to calculate class prediction scores for all tiles for all slides by performing model inference 1904, and the tiles with the highest scores from each slide (the top k tiles labeled 1906) are selected for training the model. Tiles 1906 are given the same label as the slide-level label received in 1902 (e.g., tiles for slides with class 0 are given a label of class 0). After selecting tiles from all slides, the model is trained for a single epoch (or iteration) in process 1908.
[0356] At the end of a training epoch (where the tile was used once to update the model weights), framework 1900 is again used to calculate new prediction scores for all tiles for all slides. The new prediction scores are used to identify new tiles to use for training. This process of scoring tiles and selecting the top k tiles for training is repeated until the model reaches a stopping criterion (e.g., convergence in performance on the set of held-out slides used for validation), as determined by process 1910.
[0357] Using a weakly supervised learning architecture like architecture 1800 in Figure 18 offers several advantages. Strongly supervised annotation and local annotation, such as having a pathologist manually mark examples of tissue classes across images, are costly and inefficient. Architecture 1800 can train models using a single label on datasets with orders of magnitude more slides. Furthermore, some classification targets may be impossible to annotate. For example, genetic mutations currently discovered using genotypic biomarkers may correlate with tissue features, but the identity of these tissue features has previously been unknown. However, with our technology, genotypic biomarkers can be used as slide-level labels, and the training framework of architecture 1800 can be used to identify tissue morphologies that can be used to predict the genotype present. While traditional RNA / DNA analysis to predict genotypes can take weeks, image classification to predict genotypes using our technology can be performed in a few hours and on a larger training set.
[0358] To avoid overfitting situations, in some examples, the tile selection controller 1822 can be configured to perform randomized tile selection. For example, when training with a small dataset (<300 slide images), a framework such as the framework 1900 of FIG. 19 may overfit to some tiles of class 1 slides. This situation can occur because if a class 1 tile is used for training, its score will be higher in the next epoch. Therefore, the class 1 tile will likely be selected for training again in the next epoch, further increasing its inference score and making it more likely to be selected again.
[0359] FIG. 20 illustrates a framework 2000 that can be used to avoid overfitting. Framework 2000 is similar to framework 1900 and has similar reference numbers, but uses a random tile selector 2012 to randomly select from among the high-scoring tiles, and the randomly selected tile is then sent to the training process 2008. Model 2004 is still used to determine inference scores, and tiles are selected based on the scores. However, if a tile has a high inference score, it may be considered likely to be a Class 1 tile even if it is not one of the top k tiles. For example, framework 2000 may set a low threshold score of 0.9 (or any value), and then any tile with a score of 0.9 or greater may be used as a training example. That is, any of tiles 2006 may be used for training because their scores exceed a determined threshold (e.g., a threshold of 0.9). The random high-scoring tile selector 2012 then randomly determines which of these tiles will be provided to the model training process 2008. In other examples, the tiles sent to the random high-scoring tile selector 2012 are the top k tiles. Additionally, in some examples, the tile selection probabilities applied by the selector 2012 may be completely random across all tile scores, while in other examples, the selection probabilities may be partially random, with tiles with particular scores or within particular score ranges having different random selection probabilities than tiles with other scores or within other score ranges.
[0360] FIG. 21 shows another framework 2100 for addressing overfitting situations. In some examples, when training on smaller datasets (<300 slide images), the framework of FIG. 19 may overfit slide images by predicting all tiles of class 0 slides as class 0 and all tiles of class 1 slides as class 1 (although not all tiles of class 1 slides are actually class 1). This means that tiles of class 1 slides may be misclassified. However, in framework 2100, random tile selection is performed on high-scoring tiles 2101, similar to framework 2000, but in addition, random tile selection is performed on low-scoring tiles 2103. In the illustrated example, a random low-scoring tile selector 2102 feeds a class 0 model training process 2104. A random high-scoring tile selector 2106 is provided to a class 1 model training process 2108.
[0361] The examples in Figures 19-26 are described in the context of single-class training. Label-free training of the present technology can also be used for multiclass training. For multiclass problems, a set of slide-level labels can be configured so that no slides have a class 0 label. For example, when training a model to predict consensus molecular subtypes (CMS) in colorectal cancer (CRC), CMS classes are used to guide targeted therapy, but only genotypic biomarkers, i.e., mutations in RNA data, are used. However, the present technology predicts genotypes through imaging, allowing targeted therapy to be initiated and tested in patients within hours, rather than having to wait weeks for RNA analysis. While an example is provided in reference CMS, classifier modules for other image-based biomarkers herein can be trained with architecture 1800, including PD-L1, TMB, and others.
[0362] One challenge with multiclass training is that not all images contain all classes. For example, in the case of CMS, the available labels for training images are CMS1, CMS2, CMS3, or CMS4. However, not all tiles in each image contain biomarkers indicative of one of these four classes. This can lead to situations where there are no class 0 slides that could be used to identify class 0 tiles, potentially causing the trained model to misclassify tissue types with no predictive value and reducing model accuracy. Furthermore, slides may contain features representing classes other than the slide-level label. For example, in CMS, there may be two or more subtypes of CRC. This means that a particular image may be labeled CMS1 because it is the predominant subtype in that sample, but the full-slide image may contain CMS2 tissue.
[0363] To achieve multi-class training, in some examples, the tile selection controller 1822 of FIG. 18 is configured to execute a series of processes to identify tissue features correlated with different class subtypes, e.g., the four CMS subtypes. First, the tile selection controller 1822 may be trained by identifying only positive in-class examples. Second, the tile selection controller 1822 may apply model training by identifying positive in-class tiles and negative out-of-class tiles with low positive in-class scores. Third, the tile selection controller 1822 may apply model training by identifying positive in-class tiles and negative out-of-class tiles with high negative scores. Examples of each process are now described with reference to training a CMS biomarker classifier, as shown in FIGS. 22-26.
[0364] FIG. 22 illustrates a framework 2200 that can be used to identify the CMS class with which each tissue feature most correlates, regardless of whether that tissue feature is relevant for classification. FIG. 22 illustrates an example of the first process, which identifies only examples within the positive class and shows only two classes for simplicity. For each tile in the histopathology image, a list of possible scores is shown, with each row corresponding to a class (0, 1, and 2) shown on the left. As shown, for slide images of class 1 (left side of FIG. 22), tiles with high scores for class 1 (shaded) are used for training, and for slide images of class 2 (right side of FIG. 22), tiles with high scores for class 2 (shaded) are used for training.
[0365] In this example, all tiles are classified as 1 or 2, with a 0.00 probability of being classified as class 0. Figure 23 shows the resulting overlay map illustrating the classification of CMS biomarkers with four classes. The tiles are color-coded based on the inference scores of the four CMS classes: CMS1 (microsatellite instability immune) is shown in red, CMS2 (epithelial gene expression profile, WNT and MYC signaling activation) is shown in green, CMS3 (epithelial profile with evident metabolic dysregulation) is shown in dark blue, and CMS4 (mesenchymal, pronounced transforming growth factor-β activation) is shown in light blue. In the example shown, the transparency of the tiles is adjusted to indicate the inference score, with higher scores being more opaque and lower scores being more transparent.
[0366] Figure 24 illustrates the framework 2200 applied in the second process. That is, model training continues by identifying positive in-class tiles and negative out-of-class tiles with low positive in-class scores. The first process in Figure 22 classifies all tiles as one of the non-zero classes, while the second process in Figure 24 identifies tiles that are likely background class 0. In the illustrated example, the process does this by identifying tiles with scores below the slide-level class threshold. If a tile in a class 1 slide image has a score below 0.1, the low score is marked as a slide image that can be used as an example of class 0. A similar process is performed for class 2 slide images. Figure 25 illustrates the resulting overlay map showing the CMS classification. If a tile is predicted to be class 0, it becomes transparent. Compared to Figure 23, in Figure 25, this second process identifies some tissue types as class 0. This means that the image is background tissue that does not correlate with any of the CMS classes. At the same time, the second process can identify various tissue types or tissue features associated with each of the four classes. Tiles are colored based on the inference scores of the four CMS classes as shown in Figure 23, but tiles that do not display a color are predicted to be class 0 tiles with no class prediction value.
[0367] FIG. 26 illustrates the framework 2200 applied in the third process, namely, continuing model training by identifying positive in-class tiles and negative out-of-class tiles with high negative scores. As noted above, for CMS biomarkers, the CMS classes of training images are not mutually exclusive. A CMS1 slide image may contain a CMS2 tissue type or tissue feature. Thus, in some instances, if the tile selection controller 1822 continues to label tiles as class 0 due to low slide-level class scores, some tiles may be mislabeled. From the second process above, class 0 tissues have already been identified that have low or no correlation with any of the CMS classes. Tiles with high class 0 scores exist. Thus, in the third process shown in FIG. 26, class 0 tiles may be identified based on having high class 0 scores.
[0368] Returning to FIG. 18 , architecture 1800 can be used to train any of biomarker classification models 1810-1816 to classify different biomarkers. CMS is described with reference to FIGS. 22-26 as an example. Furthermore, architecture 1800 may be agnostic to the convolutional neural network configuration, i.e., each module 1810-1816 may have the same or different configuration. In addition to the FCN architecture of FIGS. 10A-10C, modules 1810-1816 may be configured with a ResNet architecture, such as that shown in FIG. 27. The ResNet architecture provides skip connections, which help avoid the vanishing gradient problem during training as well as the degradation problem of large architectures. Pre-trained ResNet models of various sizes can be used to initialize the training of new models, including, but not limited to, ResNet-18 and ResNet-34. In addition to ResNet, architecture 1800 can be used to train neural networks with convolutional layers that create feature maps that are input into fully connected classification layers, such as AlexNet or VGG; networks with layered modules that apply multiple convolutional kernels simultaneously, such as Inception v3; networks designed to require fewer parameters for increased processing speed, such as MobileNet, SqueezeNet, and MNASNet; or customized architectures designed to better extract relevant pathological features, either manually designed or using neural architecture search networks such as NASNet.
[0369] By using a tile-based training architecture 1800 with a tile selection controller that is independent of the neural network configuration, it is possible to create a deep learning framework that can identify biomarkers where a single architecture (e.g., FCN) alone is less accurate or requires a more timely process.
[0370] Additionally, because the tile-based training architecture 1800 has a feedback configuration, in some examples, the deep learning framework 1802 can classify regions in training images, for example, using an FCN architecture, and feed those classified images back to the tile selection controller 1822 as weakly supervised training images. For example, an FCN architecture can be used to first identify specific tissue regions (e.g., tumor, stroma), which can then be used as input to a weakly supervised training pipeline such as MIL. Models trained by the weakly supervised pipeline can be used in combination with the FCN architecture to validate or improve the results of the FCN architecture or to detect new features to complement the FCN architecture.
[0371] Additionally, there exist biomarkers that are not possible to annotate, e.g., previously undiscovered tissue features that correlate with genotype, gene expression, or patient metadata. In such cases, genotype, gene expression, or patient metadata can be used to create slide-level labels that are used to train a secondary model or the FCN architecture itself, using a weakly supervised framework to discover novel classifications.
[0372] Additionally, architecture 1800 can provide for the detection of tissue and tissue artifacts. Regions within a slide image containing tissue can be first detected for use as input to the FCN architecture model. Imaging techniques such as color or texture thresholding can be used to identify tissue regions, and deep learning convolutional neural network models (e.g., FCN architectures) herein can be used to further improve the generalizability and accuracy of tissue detection. Color or texture thresholding can also be used to identify spurious artifacts within tissue images, and weakly supervised deep learning techniques can be used to improve generalizability and accuracy.
[0373] Additionally, architecture 1800 can provide marker detection. Histopathology images may contain notations or annotations drawn by a pathologist on the slide with a marker. For example, this indicates macroanatomical regions where tissue DNA / RNA analysis should be performed. A marker detection model similar to the tissue detection model can be used to identify the regions selected by the pathologist for analysis. This further complements data processing for weakly supervised training, isolating regions where DNA / RNA analysis should be performed, resulting in slide-level labels.
[0374] In FIG. 28, a process 2800 is provided for determining a suggested immunotherapy treatment for a patient using the biomarker predictions of the imaging-based biomarker prediction system 102 of FIG. 1, and in particular the deep learning framework 300 of FIG. 3. Initially, histopathology images, such as stained H&E images, are received by the system 102 (2802). In process 2804, each histopathology image is applied to a trained deep learning framework, such as one that implements one or more FCN classification configurations described herein. In process 2806, the trained deep learning framework applies the image to a trained tissue classifier model and a trained biomarker segmentation model to determine the biomarker status of tissue regions of the image. In some examples, the trained cell segmentation classifier model is further used by process 2806. In process 2806, a biomarker status and biomarker metrics for the image are generated. 29 , the output from process 2806 is provided to process 2808, which may be implemented in a tumor treatment decision system 2900 (which may be part of a genome sequencing system, an oncology system, a chemotherapy decision system, an immunotherapy decision system, or other treatment decision system) that determines the tumor type based on the received data, including based on biomarker metrics, genome sequencing data, etc. System 2900 analyzes the biomarker status and / or biomarker metrics and other received molecular data against available immunotherapies 2902 in process 2810, and system 2900 recommends a matched list of potential tumor-type-specific immunotherapies 2904, filtered from the list of available immunotherapies 2902, in the form of a corresponding treatment report.
[0375] In various examples, the imaging-based biomarker prediction system herein may be partially or entirely deployed within a dedicated slide imager, such as a high-throughput digital scanner. FIG. 30 illustrates an exemplary system 3000, which includes a dedicated ultra-high-speed pathology (slide) scanner system 3002, such as the Philips IntelliSite Pathology Solution available from Koninklijke Philips of Amsterdam, The Netherlands. In some examples, the pathology scanner system 3002 may include multiple trained biomarker classification models. Exemplary models may include, for example, models disclosed in U.S. Patent Application No. 16 / 412,362. The scanner system 3002 is coupled to an imaging-based biomarker prediction system 3004, which implements the processes described and illustrated in the examples herein. For example, in the illustrated example, system 3004 includes a deep learning framework 3006 based on tile-based multi-scale and / or single-scale classification modules, according to examples herein, having one or more trained biomarker classifiers 3008, trained cell classifiers 3010, and trained tissue classifiers 3012. Deep learning framework 3006 performs biomarker and tumor classification on the histopathology images and stores the classification data as overlay data with the original images in generated image database 3014. For example, the images may be stored as TIFF files. However, database 3014 may include JSON files and other data generated by the classification process herein. In some examples, the deep learning framework may be integrated, in whole or in part, within scanner 3002, as shown in optional block 3015.
[0376] To manage the generated images, which can be quite large, an image management system and viewer generator 3016 is provided. In the illustrated example, system 3016 is shown as being external to imaging-based biomarker prediction system 3004, connected by a private or public network. However, in other examples, all or part of system 3016 can be deployed on system 3004, as shown at 3019. In some examples, system 3016 is cloud-based and stores the generated images from (or instead of) database 3014. In some examples, system 3016 generates a web-accessible, cloud-based viewer that allows a pathologist to access, view, and manipulate histopathology images with various classification overlays via a graphic user interface. Examples are shown in FIGS. 31-37.
[0377] In some examples, the image management system 3016 manages the receipt of scanned slide images 3018 from the scanner 3002 , which are generated from the imager 3020 .
[0378] In the illustrated example, the image management system 3016 generates an executable viewer app 3024 and deploys the app 3024 to an app deployment engine 3022 of the scanner 3002. The app deployment engine 3022 can provide functionality such as GUI generation that allows a user to interact with the viewer app 3024, an app marketplace that allows a user to download the viewer app 3024 from the image management system 3016 or other network-accessible sources, and the like.
[0379] 31-37 show, in one example, various digital screenshots generated by the embedded viewer 3024, which are displayed in a GUI format that allows the user to manipulate the displayed images to zoom in and out and to view various classifications of tissues, cells, biomarkers, and / or tumors.
[0380] Referring to FIG. 31 , a GUI-generated display 3100 is shown with a panel 3102 showing an entire histopathology image 3104 and a magnified portion of that image 3104 (at 1.3x magnification) displayed as a window 3106. The panel 3102 further includes a magnification corresponding to the window 3106 and a tumor content report. FIG. 32 shows the display 3100 at 3.0x magnification after the user zooms in on the window 3106. FIG. 33 is similar, but at 5.7x magnification. FIG. 34 shows a drop-down menu 3108 listing a series of classifications that the user can select to generate a classification overlay map that is displayed on the display 3100. FIG. 35 shows the resulting display 3100 with an overlay map showing tumor-classified tissue, in this example showing the tissue divided into tiles and the tiles with classifications. In the example of FIG. 35 , the classification shown is a tumor classification. Figure 36 shows another exemplary classification overlay mapping, which is one of cell classification, epithelial, immune, stromal, tumor, or other. Figure 37 shows a magnified cell classification overlay mapping showing the classifications that can actually be displayed at sufficient magnification to distinguish different cells in a histopathology image.
[0381] FIG. 38 illustrates an exemplary computing device 3800 for implementing the imaging-based biomarker prediction system 100 of FIG. 1. As illustrated, system 100 may be implemented on computing device 3800, particularly on one or more processing units 3810, which may represent central processing units (CPUs), and / or on one or more graphics processing units (GPUs) 3811, including clusters of CPUs and / or GPUs, and / or on one or more tensor processing units (TPUs) (also labeled 3811), any of which may be cloud-based. Features and functionality described for system 100 may be stored on and implemented from one or more non-transitory computer-readable media 3812 of computing device 3800. Computer-readable media 3812 may include, for example, an operating system 3814 and a deep learning framework 3816 having elements corresponding to elements of deep learning framework 300, including pre-processing controller 302, classifier modules 304 and 306, and post-processing controller 308. More generally, the computer-readable medium 3812 can store trained deep learning models, executable code, etc. used to implement the techniques herein. The computer-readable medium 3812 and processing unit 3810 and TPU / GPU 3811 can store image data, tissue classification data, cell segmentation data, leukocyte segmentation data, TIL metrics, and other data in one or more databases 3813 herein. The computing device 3800 includes a network interface 3824 communicatively coupled to a network 3850 for communication to and / or from a portable personal computer, smartphone, electronic document, tablet, and / or desktop personal computer, or other computing device. The computing device further includes an I / O interface 3826 connected to devices such as a digital display 3828, a user input device 3830, etc.In some examples, as described herein, computing device 3800 generates biomarker predictions as electronic documents 3815 that can be accessed and / or shared over network 3850. In the illustrated example, system 100 is implemented on a single server 3800. However, the functionality of system 100 may be implemented across distributed devices 3800, 3802, 3804, etc., connected to each other via communication links. In other examples, the functionality of system 100 may be distributed across any number of devices, including the illustrated portable personal computers, smartphones, electronic documents, tablets, and desktop personal computing devices. In other examples, the functionality of system 100 may be cloud-based, such as one or more connected cloud TPUs customized to perform machine learning processes. Network 3850 may be a public network such as the Internet, a private network such as a research institution's or a company's private network, or any combination thereof. The network may include a local area network (LAN), a wide area network (WAN), cellular, satellite, or other network infrastructure, whether wireless or wired. The network may utilize communication protocols including packet-based and / or datagram-based protocols such as Internet Protocol (IP), Transmission Control Protocol (TCP), User Datagram Protocol (UDP), or other types of protocols. Additionally, the network may include multiple devices that facilitate network communications and / or form the hardware infrastructure of the network, such as switches, routers, gateways, access points (such as wireless access points as shown), firewalls, base stations, repeaters, backbone devices, etc.
[0382] The computer-readable medium may include executable computer-readable code stored on a computer (including, for example, a processor and a GPU) for programming the computer to implement the techniques herein. Examples of such computer-readable storage media include hard disks, CD-ROMs, digital versatile disks (DVDs), optical storage devices, magnetic storage devices, ROMs (read-only memories), PROMs (programmable read-only memories), EPROMs (erasable programmable read-only memories), EEPROMs (electrically erasable programmable read-only memories), and flash memories. More generally, the processing unit of the computing device 1300 may represent a CPU-type processing unit, a GPU-type processing unit, a TPU-type processing unit, a field-programmable gate array (FPGA), another class of digital signal processor (DSP), or other hardware logic components that can be driven by a CPU.
[0383] While the exemplary deep learning framework herein has been described as being configured with an exemplary machine learning architecture (FCN configuration), it should be noted that any number of suitable convolutional neural network architectures can be used. Generally speaking, the deep learning framework herein can implement any suitable statistical model (e.g., a neural network or other model implemented through a machine learning process) that is applied to each received image. As discussed herein, the statistical model can be implemented in a wide variety of ways. In some examples, machine learning is used to evaluate training images and develop a classifier that correlates predefined image features to specific categories of TIL status. In some examples, image features can be identified using a learning algorithm, such as a neural network, a support vector machine (SVM), or other machine learning process, to train the classifier. Once the classifier in the statistical model is properly trained on a set of training images, the statistical model can be used in real time to analyze subsequent images that are provided as input to the statistical model to predict biomarker status. In some examples, when the statistical model is implemented using a neural network, the neural network can be configured in a wide variety of ways. In some examples, the neural network may be a deep neural network and / or a convolutional neural network. In some examples, the neural network may be a distributed and scalable neural network. The neural network may be customized in a variety of ways, such as providing a specific top layer, such as a logistic regression top layer. A convolutional neural network may be viewed as a neural network that includes a set of nodes with associated parameters. A deep convolutional neural network may be viewed as having a structure in which multiple layers are stacked. Neural networks or other machine learning processes may include various sizes, numbers of layers, and levels of connectivity. Some layers may correspond to stacked convolutional layers (optionally followed by contrast normalization and max pooling) followed by one or more fully connected layers.For neural networks trained with large datasets, the number of layers and layer size can be increased by using dropout to address the potential problem of overfitting. In some cases, neural networks can be designed to forgo the use of fully connected upper layers at the top of the network. By forcing the network to reduce the dimensionality of intermediate layers, very deep neural network models can be designed while dramatically reducing the number of learned parameters.
[0384] A system for performing the methods described herein may include a computing device, and more specifically, may be implemented in one or more processing units, such as a central processing unit (CPU) and / or one or more graphics processing units (GPUs), including clusters of CPUs and / or GPUs. The described features and functions may be stored on and implemented from one or more non-transitory computer-readable media of the computing device. The computer-readable media may include, for example, an operating system and software modules, or "engines," that implement the methods described herein. More generally, the computer-readable media may store batch normalization process instructions for the engine to implement the techniques described herein. The computing device may be a distributed computing system, such as an Amazon Web Services cloud computing solution.
[0385] The computing device includes a network interface communicatively coupled to a network for communicating to and / or from a portable personal computer, a smartphone, an electronic document reader, a tablet, and / or a desktop personal computer, or other computing device. The computing device further includes an I / O interface connected to devices such as a digital display, a user input device, etc.
[0386] The engine's functionality may be implemented on distributed computing devices interconnected via communications links, or the like. In other examples, the system's functionality may be distributed across any number of devices, including the illustrated portable personal computers, smartphones, electronic documents, tablets, and desktop personal computing devices. The computing devices may be communicatively coupled to a network and another network. The network may be a public network such as the Internet, a private network such as a research or corporate network, or any combination thereof. The network may include a local area network (LAN), a wide area network (WAN), cellular, satellite, or other network infrastructure, whether wireless or wired. The network may utilize communications protocols including packet-based and / or datagram-based protocols, such as Internet Protocol (IP), Transmission Control Protocol (TCP), User Datagram Protocol (UDP), or other types of protocols. Furthermore, the network may include multiple devices that facilitate network communication and / or form the hardware infrastructure of the network, such as switches, routers, gateways, access points (such as wireless access points as illustrated), firewalls, base stations, repeaters, backbone devices, etc.
[0387] The computer-readable medium may include executable computer-readable code stored on a computer (including, for example, a processor and a GPU) for programming the computer to the techniques herein. Examples of such computer-readable storage media include hard disks, CD-ROMs, digital versatile disks (DVDs), optical storage devices, magnetic storage devices, ROMs (read-only memories), PROMs (programmable read-only memories), EPROMs (erasable programmable read-only memories), EEPROMs (electrically erasable programmable read-only memories), and flash memories. More generally, the processing unit of a computing device may represent a CPU-type processing unit, a GPU-type processing unit, a field programmable gate array (FPGA), another class of digital signal processor (DSP), or other hardware logic components that can be driven by a CPU.
[0388] Throughout this specification, multiple instances may implement components, operations, or structures described as a single instance. Although individual operations of one or more methods are illustrated and described as separate operations, one or more of the individual operations may be performed simultaneously, and the operations need not be performed in the order illustrated. Structures and functions presented as separate components in example configurations may be implemented as a combined structure or component. Similarly, structures and functions presented as a single component may be implemented as a separate component or multiple components.
[0389] Additionally, certain embodiments are described herein as including logic or a number of routines, subroutines, applications, or instructions. These can constitute either software (e.g., code embodied on a machine-readable medium or in a transmission signal) or hardware. In hardware, routines, etc., are tangible units capable of performing particular operations and can be configured or arranged in a particular way. In exemplary embodiments, one or more computer systems (e.g., standalone, client, or server computer systems), or one or more hardware modules of a computer system (e.g., a processor or processors), can be configured as hardware modules that operate by software (e.g., an application or portion of an application) to perform particular operations described herein.
[0390] In various embodiments, a hardware module can be implemented mechanically or electronically. For example, a hardware module can include dedicated circuitry or logic (e.g., a microcontroller, a special-purpose processor such as a field programmable gate array (FPGA), or an application-specific integrated circuit (ASIC)) that is permanently configured to perform specific operations. A hardware module can also include programmable logic or circuitry (e.g., contained within a general-purpose processor or other programmable processor) that is temporarily configured by software to perform specific operations. It will be appreciated that the decision as to whether a hardware module is implemented mechanically, with dedicated, permanently configured circuitry, or with temporarily configured circuitry (e.g., configured by software) can be determined based on cost and time considerations.
[0391] Thus, the term "hardware module" should be understood to encompass a tangible entity, that is physically constructed, permanently configured (e.g., hardwired), or temporarily configured (e.g., programmed) to operate in a particular manner or to perform certain operations described herein. Considering embodiments in which the hardware modules are temporarily configured (e.g., programmed), each of the hardware modules need not be configured or instantiated at any one instance of time. For example, if the hardware modules include a general-purpose processor configured using software, the general-purpose processor can be configured as different hardware modules at different times. Thus, the software may configure the processor, for example, to configure a particular hardware module at one time and a different hardware module at another time.
[0392] Hardware modules can provide information to and receive information from other hardware modules. Thus, the described hardware modules can be considered communicatively coupled. When multiple such hardware modules are present simultaneously, communication can be achieved via signal transmission (e.g., via appropriate circuits and buses) connecting the hardware modules. In embodiments in which multiple hardware modules are configured or instantiated at different times, communication between such hardware modules can be achieved, for example, via storage and retrieval of information in memory structures accessed by the multiple hardware modules. For example, one hardware module can perform an operation and store the output of that operation in a memory device to which the hardware module is communicatively coupled. An additional hardware module can then later access the memory device to retrieve and process the stored output. Hardware modules can also initiate communication with input or output devices to operate on resources (e.g., collect information).
[0393] Various operations of the example methods described herein may be performed, at least in part, by one or more processors that are temporarily (e.g., by software) or permanently configured to perform the associated operations. Whether temporarily or permanently configured, such processors may constitute processor-implemented modules that operate to perform one or more operations or functions. Modules referred to herein may, in some example embodiments, include processor-implemented modules.
[0394] Similarly, methods or routines described herein may be at least partially processor-implemented. For example, at least some of the operations of a method may be performed by one or more processors or processor-implemented hardware modules. Certain performance of an operation may reside not only within a single machine, but also be distributed among one or more processors deployed across several machines. In some embodiments, one or more processors may reside in a single location (e.g., in a home environment, in a work environment, or as a server farm), while in other embodiments, the processors may be distributed across multiple locations.
[0395] Certain performance aspects of an operation may reside not only within a single machine but may also be distributed among one or more processors deployed across several machines. In some exemplary embodiments, one or more processors or processor-implemented modules may reside in a single location (e.g., in a home environment, in a work environment, or as a server farm). In other exemplary embodiments, one or more processors or processor-implemented modules may be distributed across multiple locations.
[0396] Unless otherwise indicated, descriptions herein using words such as "processing," "computing," "calculating," "determining," "presenting," and "displaying" may refer to machine (e.g., computer) operations or processes that manipulate or transform data represented as physical (e.g., electronic, magnetic, or optical) quantities in one or more memories (e.g., volatile memory, non-volatile memory, or a combination thereof), registers, or other machine components that receive, store, transmit, or display information.
[0397] As used herein, any reference to "one embodiment" or "an embodiment" means that a particular element, feature, structure, or characteristic described in connection with an embodiment is included in at least one embodiment. The appearances of the phrase "in one embodiment" in various places in the specification are not necessarily all referring to the same embodiment.
[0398] Some embodiments may be described using the terms "coupled" and "connected," along with their derivatives. For example, some embodiments may be described using the term "coupled" to indicate that two or more elements are in direct physical or electrical contact. However, the term "coupled" can also mean that two or more elements are not in direct contact with each other, but yet still cooperate or interact with each other. The embodiments are not limited in this context.
[0399] As used herein, the terms "comprises," "comprising," "includes," "including," "has," or any other variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, article, or apparatus that includes a list of elements is not necessarily limited to only those elements and may include other elements not expressly listed or inherent to such process, method, article, or apparatus. Furthermore, unless expressly stated to the contrary, "or" means an inclusive or, not an exclusive or. For example, condition A or B is satisfied by any one of A being true (or present) and B being false (or absent), A being false (or absent) and B being true (or present), and both A and B being true (or present).
[0400] Additionally, the use of "a" or "an" is used to describe elements and components of the embodiments herein. This is done merely for convenience and to give a general sense of the description. This description should be read to include one or at least one, and the singular also includes the plural unless it is clear that otherwise is meant.
[0401] This detailed description should be construed as merely exemplary and does not describe every possible embodiment, as describing every possible embodiment would be impractical, if not impossible. Many alternative embodiments may be implemented using either current technology or technology developed after the filing date of this application.
[0402] [1] 1. A computer-implemented method for identifying biomarkers in digital images of hematoxylin and eosin (H&E) stained slides of a target tissue, the method comprising: receiving the digital image into an image-based biomarker prediction system having one or more processors; using the one or more processors to perform an image tiling process on the digital image by separating the digital image into a plurality of tile images, each of the plurality of tile images comprising a different portion of the digital image; using the one or more processors to apply the plurality of tile images to a multi-scale deep learning framework including one or more trained deep learning multi-scale classifier models, each trained to classify a different tissue classification for each tile image, and using the multi-scale deep learning framework to determine a tissue classification for each of the plurality of tile images; using the one or more processors to identify cells in the digital image using a trained cell segmentation model; and identifying a predicted presence of one or more biomarkers associated with the digital image from the tissue classification determined for each tile image and from the identified cells within the digital image. [2] Item 10. The method of item 1, wherein performing the image tiling process on the digital image comprises applying a tiling mask to the digital image to separate the digital image into the plurality of tile images. [3] Item 3. The method of item 2, wherein the tiling mask comprises tiles of the same size. [4] Item 3. The method of item 2, wherein the tiling mask comprises tiles of different sizes. [5] Item 3. The method of item 2, wherein the tiling mask comprises tiles having a rectangular shape. [6] Item 3. The method of item 2, wherein the tiling mask comprises tiles characterized by the topology and / or morphology of pixels or groups of pixels. [7] receiving the digital image using the one or more processors; acquiring the digital image at a first image resolution; downsampling the digital image to a second image resolution; performing luminance normalization on the pixels of the digital image; and removing non-tissue objects from the digital image. [8] the one or more trained deep learning multi-scale classifier models; receiving, in the multi-scale deep learning framework, a plurality of H&E slide training images from a training image dataset, each H&E slide training image having a label corresponding to a biomarker to be trained; performing a tile-based tissue classification analysis on each of the H&E slide training images; performing a pixel-based cell segmentation analysis on each of the H&E slide training images; Optionally, performing a tile-based biomarker classification analysis on each of said H&E slide training images; 2. The method of claim 1, further comprising generating the one or more trained deep learning multi-scale classifier models accordingly. [9] 9. The method of claim 8, wherein each H&E slide training image includes multiple tile images, each having a tile-level label.
[10] 9. The method of claim 8, further comprising, for each H&E slide training image, labeling each of a plurality of tile images of the H&E slide training image with a tile level label.
[11] performing, for each H&E slide training image, a tile selection process that infers a class status for each tile image within the H&E slide training image; 9. The method of claim 8, further comprising: discarding tile images that do not correspond to a class of interest before performing the tile-based tissue classification analysis on each of the H&E slide training images based on the inferred class status, thereby performing the tile-based tissue classification analysis only on selected tile images of the H&E slide training images.
[12] the one or more trained deep learning multi-scale classifier models; receiving a molecular training dataset of a plurality of training tissue samples, the molecular training dataset including RNA transcriptome counts from sequencing of substantially similar samples associated with each training tissue sample; performing a clustering process on the molecular training dataset to identify one or more molecular data subsets, each corresponding to a different biomarker; for each of the one or more molecular data subsets, identifying a plurality of digital images of H&E stained training slides of training tissue samples corresponding to the respective biomarkers for an image-based biomarker prediction system having one or more processors; performing a tile-based tissue classification analysis on each of the H&E slide training images; performing a pixel-based cell segmentation analysis on each of the H&E slide training images; Optionally, performing a tile-based biomarker classification analysis on each of said H&E slide training images; 2. The method of claim 1, further comprising generating the one or more trained deep learning multi-scale classifier models accordingly.
[13] Item 1. The method of item 1, wherein each of the one or more trained deep learning multi-scale classifier models is configured as a tile-resolution fully convolutional network (FCN) classification model.
[14] Item 1. The method of item 1, wherein the trained deep learning multi-scale classifier model is trained to classify tissues as corresponding to a classification selected from the group consisting of tumor, stromal, normal, lymphocyte, fat, muscle, blood vessel, immune cluster, necrosis, hyperplasia / dysplasia, and red blood cell, respectively.
[15] identifying cells within the digital image tiles using the trained cell segmentation model; 2. The method of claim 1, further comprising using the one or more processors to apply each of the plurality of tile images to the cell segmentation model and, for each tile, assigning a cell classification to one or more pixels in the tile image.
[16] assigning the cell classification to one or more pixels in the tile image; 16. The method of claim 15, comprising using the one or more processors to identify the one or more pixels as a cell interior, a cell boundary, or a cell exterior, and classifying the one or more pixels as the cell interior, the cell boundary, or the cell exterior.
[17] Item 10. The method of item 1, wherein the trained cell segmentation model is a pixel-resolution three-dimensional UNet classification model trained to classify cell interiors, cell boundaries, and cell exteriors.
[18] identifying cells within the digital image tiles using the trained cell segmentation model; applying, using the one or more processors, each of the plurality of tile images to the cell segmentation model; and performing alignment of each segmented cell of the tile image by using the one or more processors to determine a cell boundary of each cell, determine a center of gravity of each cell, and shift the coordinates of the center of gravity into a universal coordinate space of the digital image.
[19] Item 10. The method of item 1, wherein the trained cell segmentation model is trained using a set of annotated H&E slide training images that identify cell boundaries, cell interiors, and cell exteriors.
[20] 2. The method of claim 1, wherein the digital image is an unlabeled digital image or a slide-level labeled image. [twenty one] Item 10. The method of item 1, wherein the digital image is a tile-level labeled image. [twenty two] 2. The method of item 1, wherein the one or more biomarkers are selected from the group consisting of tumor-infiltrating lymphocytes (TIL), nuclear-to-cytoplasmic (NC) ratio, ploidy, signet ring morphology, and programmed death-ligand 1 (PD-L1). [twenty three] wherein the one or more biomarkers are TILs, and the method comprises: 2. The method of claim 1, further comprising using the one or more processors to identify lymphocyte cells in the digital image using the lymphocyte segmentation model by integrating, using the one or more processors, the cell boundaries identified using the cell segmentation model with the lymphocytes identified using a trained lymphocyte segmentation model, and using the one or more processors to generate a nested classification of each cell. [twenty four] wherein the one or more biomarkers are TILs, and the method comprises: 2. The method of claim 1, further comprising using the one or more processors to identify lymphocyte cells in the digital image using a trained lymphocyte segmentation model by applying each of a plurality of tile images to a trained lymphocyte segmentation model and, for each tile image, assigning a lymphocyte classification to one or more pixels in the tile image. [twenty five] 2. The method of claim 1, wherein the one or more biomarkers are TILs, and the method further comprises using the one or more processors to identify lymphocyte cells in the digital image using a trained lymphocyte segmentation model, wherein the lymphocyte segmentation model is a pixel-resolution two-dimensional UNet classification model trained to distinguish between lymphocyte cell classifications within cell boundaries and non-lymphocyte cell classifications within cell boundaries.
[26] Item 10. The method of item 1, wherein the one or more processors are one or more graphics processing units (GPUs), tensor processing units (TPUs), and / or central processing units (CPUs).
[27] 2. The method of claim 1, wherein the image-based biomarker prediction system is communicatively coupled to a pathology slide scanner system via a communication network, whereby the image-based biomarker prediction system receives the digital images from the pathology slide scanner system via the communication network.
[28] 2. The method of claim 1, wherein the image-based biomarker prediction system is comprised within a pathology slide scanner system.
[29] 2. The method of claim 1, wherein at least one of the one or more processors of the image-based biomarker prediction system is included within a pathology slide scanner system.
[30] 1. A computer-implemented method for identifying biomarkers in digital images of hematoxylin and eosin (H&E) stained slides of a target tissue, the method comprising: receiving a molecular training dataset of a plurality of training tissue samples, the molecular training dataset including RNA transcriptome counts from sequencing of substantially similar samples associated with each training tissue sample; performing a clustering process on the molecular training dataset to identify one or more molecular data subsets, each corresponding to a different biomarker; for each of the one or more molecular data subsets, receiving a plurality of digital images of H&E stained training slides of training tissue samples corresponding to the respective biomarkers for an image-based biomarker prediction system having one or more processors; generating, using the one or more processors, for each of the one or more molecular data subsets, a trained image-based biomarker classifier model based on the plurality of digital images of the H&E stained training slides; receiving, using said one or more processors, a subsequent digital image of an H&E stained slide of the subsequent tissue sample; and applying, using the one or more processors, the subsequent digital image to the trained image-based biomarker classification model to identify a predicted presence of one or more biomarkers in the subsequent tissue sample.
[31] 31. The method of claim 30, wherein generating the trained image-based biomarker classifier model for each of the one or more molecular data subsets comprises performing a multi-instance learnin...
Claims
1. 1. A method for digital image analysis of a slide containing tissue for predicting PD-L1 status, comprising: The method comprises the steps of: dividing the digital image of the slide into a plurality of tiles, each tile of the plurality of tiles including a corresponding portion of the digital image; For each tile of the plurality of tiles: identifying a set of characteristics of the tile; identifying a set of structural texture features in a second portion of the digital image that includes at least a portion of at least one additional tile other than the tile, wherein the second portion of the digital image is larger than an area of the digital image that corresponds to the tile; and determining initial predictive PD-L1 metrics for the tiles based on the set of tile features and the set of structural texture features of the second portion of the digital image; Including; identifying a plurality of cellular objects within the digital image; determining a subsequent predicted PD-L1 metric for each cell object of said plurality of cell objects; for each tile of the plurality of tiles that corresponds to one of the plurality of cell objects, replacing the corresponding initial predicted PD-L1 metric with the subsequent predicted PD-L1 metric for that cell object; and For each tile of the plurality of tiles corresponding to one of the plurality of cell objects, determining a predicted PD-L1 status of tissue contained on the slide based on replacing the initial predicted PD-L1 metrics with subsequent predicted PD-L1 metrics. The method comprising:
2. The step of identifying a set of tile characteristics includes: The method of claim 1 , comprising identifying at least one of color, texture, cell size, shape, and spatial organization.
3. The step of identifying a set of structural texture characteristics of a second portion of the digital image includes: The method of claim 1 , comprising identifying at least one of glands, ducts, blood vessels, and immune clusters.
4. The method of claim 1 , wherein the slide is stained with hematoxylin and eosin (H&E).
5. The method of claim 1 , wherein identifying the set of tile features comprises identifying the set of tile features using a trained deep learning model.
6. The method of claim 5 , wherein the trained deep learning model is configured as a tile-resolution fully convolutional network (FCN) classification model.
7. The method of claim 1 , wherein identifying the set of structural tissue features comprises identifying the set of structural tissue features using a trained deep learning model.
8. The method of claim 7 , wherein the trained deep learning model is configured as a tile-resolution fully convolutional network (FCN) classification model.
9. The step of identifying the plurality of cell objects includes the steps of: applying each tile of the plurality of tiles to a cell segmentation model; and 2. The method of claim 1, further comprising: determining a cell boundary and a center point of each segmented cell in each tile of the plurality of tiles, and performing registration by shifting the coordinates of the center point into a universal coordinate space of the digital image.
10. 10. The method of claim 9, wherein the cell segmentation model is trained using a set of H&E slide training images in which cell boundaries, cell interiors, and cell exteriors are annotated.
11. The method of claim 1 , wherein the plurality of cellular objects comprises at least one of CD3, CD8, CD20, pancytokeratin, and smooth muscle actin.
12. The method of claim 1 , wherein the plurality of cellular objects includes lymphocytes and non-lymphocytes.
13. 1. A computing device configured to identify a predicted PD-L1 status of a digital image of a slide, the computing device comprising: one or more memories; and Interfaced with one or more memories to perform the following operations: dividing the digital image of the slide into a plurality of tiles, each tile of the plurality of tiles including a corresponding portion of the digital image; For each tile of the plurality of tiles: identifying a set of characteristics of the tile; identifying a set of structural texture features in a second portion of the digital image that includes at least a portion of at least one additional tile other than the tile, wherein the second portion of the digital image is larger than an area of the digital image that corresponds to the tile; and determining initial predictive PD-L1 metrics for the tiles based on the set of tile features and the set of structural texture features of the second portion of the digital image; Including; identifying a plurality of cellular objects within the digital image; determining a subsequent predicted PD-L1 metric for each cell object of said plurality of cell objects; for each tile of the plurality of tiles that corresponds to one of the plurality of cell objects, replacing the corresponding initial predicted PD-L1 metric with the subsequent predicted PD-L1 metric for that cell object; and For each tile of the plurality of tiles corresponding to one of the plurality of cell objects, determining a predicted PD-L1 status of tissue contained on the slide based on replacing the initial predicted PD-L1 metrics with subsequent predicted PD-L1 metrics.
10. The computing device, comprising: one or more processors configured to:
14. The computing device of claim 13 , wherein at least one of the one or more processors is included within a pathology slide scanner system.
15. A non-transitory computer-readable storage medium having stored thereon a set of instructions executable by a processor, the instructions including: instructions for dividing a slide of the digital image into a plurality of tiles, each tile of the plurality of tiles constituting a corresponding portion of the digital image; For each tile of the plurality of tiles: Identifying a set of tile features; identifying a set of structural texture features of a second portion of the digital image that includes at least a portion of at least one additional tile of the plurality of tiles other than the tile, wherein the second region of the digital image is larger than the region of the digital image that corresponds to the tile; and determining an initial predicted PD-L1 metric for the tile based on the set of features of the tile and the set of structural organization features of a second region of the digital image; and instructions for identifying a plurality of cellular objects within the digital image; instructions for performing a subsequent determination of a predictive PD-L1 metric for each cell object of the plurality of cell objects; instructions for, for a tile of the plurality of tiles corresponding to one of the plurality of cell objects, replacing the corresponding initial predicted PD-L1 metric with the subsequent predicted PD-L1 metric for the cell object; and and instructions for determining a predicted PD-L1 status of tissue contained on the slide based on replacing an initial predicted PD-L1 metric with a subsequent predicted PD-L1 metric for each tile of a plurality of tiles corresponding to one of a plurality of cell objects corresponding to the plurality of tiles. The medium comprising:
Citation Information
Patent Citations
Systems and methods for detection of biological structures and / or patterns in images
JP2017516992A
Methods and systems for determining ratios of different cell subsets
JP2018512071A