Method for determining biomarker from pathological tissue slide image
A deep learning framework efficiently processes large digital images of tumor samples by using multi-scale and single-scale configurations to identify biomarkers like TILs and PD-L1, addressing inefficiencies in existing methods and improving cancer diagnosis and treatment planning.
Patent Information
- Application Number
- JP2025081459
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2020-02-28
- Filing Date
- 2025-05-14
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2040-03-25
AI Technical Summary
Current methods for analyzing tumor samples, such as conventional CNNs and FCNs, are inefficient in processing large digital whole-slide images due to redundant computation and long execution times, making it difficult to accurately diagnose biomarkers like TILs and PD-L1 in cancer tissues.
A deep learning framework is developed to analyze pathological tissue images, utilizing multi-scale and single-scale configurations with tile-level and pixel-level classifiers, enabling efficient identification of biomarkers like TILs and PD-L1 by separating images into tiles and applying trained classifiers, reducing redundant computation and improving processing time.
The deep learning framework efficiently identifies biomarkers in large digital images, facilitating optimized drug treatment recommendations and improving disease progression prediction, thereby enhancing cancer diagnosis and treatment planning.
Smart Images

Figure 2025113294000001_ABST
Abstract
Description
Technical Field
[0001] Cross - Reference to Related Applications This application is a continuation - in - part of U.S. Patent Application No. 16 / 732,242, filed on December 31, 2019, which claims the priority of U.S. Provisional Patent Application No. 62 / 787,047, filed on December 31, 2018, and is also a continuation - in - part of U.S. Patent Application No. 16 / 412,362, filed on May 14, 2019, which claims the priority of U.S. Provisional Patent Application No. 62 / 671,300, filed on May 14, 2018, and claims the priority of U.S. Provisional Patent Application No. 62 / 824,039, filed on March 26, 2019, U.S. Provisional Patent Application No. 62 / 889,521, filed on August 20, 2019, and U.S. Provisional Patent Application No. 62 / 983,524, filed on February 28, 2020. The entire disclosure of each of them is hereby expressly incorporated by reference herein.
[0002] The present disclosure relates to detecting, quantifying, and / or characterizing biomarkers associated with cancer, and more particularly to examining digital images for detecting, quantifying, and / or characterizing such biomarkers from the analysis of one or more histological tissue slide images.
Background Art
[0003] The description of the background art provided herein is for the purpose of generally presenting the context of the present disclosure. The research of the inventors named as such at present, to the extent described in this background art section, and aspects of the description that may not be considered prior art at the time of filing in another way are not admitted as prior art to the present disclosure, either expressly or implicitly.
[0004] To guide medical experts in the diagnosis, prognosis, and treatment evaluation of a patient's cancer, it is common to extract and examine tumor samples from the patient. Visual inspection can reveal the growth pattern of cancer cells within the tumor and the presence of immune cells within the tumor in relation to healthy cells near the cancer cells. Conventionally, a thin slice of tumor tissue mounted on a glass microscope slide has been visually analyzed by a pathologist, a member of a pathology team, other trained medical experts, or other human analysts to identify each region of tissue corresponding to one of the many tissue types present in the tumor sample. This information helps the pathologist determine the characteristics of the patient's cancer tumor and may be useful in making treatment decisions. Pathologists often assign one or more numerical scores to the slide based on visual approximation.
[0005] To make these visual approximations, medical experts have attempted to identify many characteristics of the tumor, including, for example, the malignancy of the tumor, the purity of the tumor, the degree of tumor invasiveness, the degree of immune infiltration into the tumor, the stage of the cancer, and the anatomical origin site of the tumor, which can be important in the diagnosis and treatment of metastatic tumors. These details about the cancer can help physicians monitor the progression of the cancer within the patient and predict which anticancer treatments are likely to be successful in eliminating cancer cells from the patient's body.
[0006] As another feature of tumors, there may be specific biomarkers or other cell types, including immune cells, within or near the tumor. For example, tumor-infiltrating lymphocytes (TILs) present at high levels are recognized as biomarkers of anti-tumor immune responses across a wide range of tumors. TILs are mononuclear immune cells that infiltrate tumor tissue or stroma and have been reported in multiple tumor types, including breast cancer. The population of TILs is composed of various types of cells (i.e., T cells, B cells, natural killer (NK) cells, etc.). The population of TILs that occurs naturally in cancer patients is rarely effective in destroying tumors, but the presence of TILs has been associated with improved prognosis in many types of cancer, such as epithelial ovarian cancer, colon cancer, esophageal cancer, melanoma, endometrial cancer, and breast cancer (see, for example, Melichar et al., Anticancer Res. 2014;34(3):1115-25, Naito et al., Cancer Res. 1998;58(16):3491-4).
[0007] Another characteristic of tumors is the presence of specific molecules as biomarkers, which includes a molecule known as programmed death ligand 1 (PD-L1). PD-L1 is associated with the diagnosis and evaluation of non-small cell lung cancer (NSCLC). NSCLC affects over 1.5 million people worldwide and is the most common type of lung cancer. NSCLC has a poor response to standard chemoradiation therapy and a high recurrence rate, resulting in a low 5-year survival rate. With the advancement of immunology, it has been found that in NSCLC, the expression of PD-L1, which binds to programmed death-1 (PD-1) expressed on the surface of T cells, frequently increases. The binding of PD-1 and PD-L1 inactivates the anti-tumor response of T cells, enabling the avoidance of targeting by the immune system in NSCLC. The discovery of the interaction between tumor progression and the immune response has led to the development and regulatory approval of PD-1 / PD-L1 checkpoint blockade immunotherapies such as nivolumab and pembrolizumab. Anti-PD-1 and anti-PD-L1 antibodies restore the anti-tumor immune response by interfering with the interaction between PD-1 and PD-L1. In particular, in PD-L1-positive NSCLC patients treated with these checkpoint inhibitors, sustained tumor regression and improved survival have been achieved.
[0008] As the role of immunotherapy in oncology expands, accurately assessing the PD-L1 status of tumors can help identify patients in whom PD-1 / PD-L1 checkpoint blockade immunotherapy may be effective. Currently, immunohistochemical (IHC) staining of tumor tissue obtained from biopsies or surgical specimens is employed to evaluate the PD-L1 status. However, such IHC staining is often limited due to insufficient tissue samples or lack of resources in some settings.
[0009] Hematoxylin and eosin (H&E) staining is a long-established method used by pathologists to analyze the morphological characteristics of tissues for the diagnosis of malignant tumors. For example, H&E slides can show visual features of tissue structures such as cell nuclei and cytoplasm, which can be useful for the identification of cancerous tumors.
[0010] With the advancement of technology, it has become possible to digitize histopathological H&E and IHC slides into high-resolution whole-slide images (WSIs), providing an opportunity to develop computer vision tools for a wide range of clinical applications. By converting microscope slides into high-resolution digital images, computer-aided slide analysis becomes possible, enabling the classification of tissue types or pathological classifications. Generally speaking, for example, deep learning applications have shown promise as tools in medical diagnosis applications and the prediction of treatment outcomes. Deep learning is a subset of machine learning, and models can be constructed with multiple individual neural node layers. A convolutional neural network ("CNN") is a neural network that employs convolutional techniques. For example, a CNN can provide a deep learning process by which digital images are analyzed by assigning one class label to each input image. However, WSIs contain two or more types of tissue, including the boundaries between adjacent tissue classes. In order to analyze the boundaries between adjacent tissue classes and the presence of immune cells among tumor cells, it is necessary to classify partially different regions as different tissue classes. Since traditional CNNs assign multiple tissue classes to a single slide image, the CNN needs to process each section of the image that requires the assignment of tissue class labels individually. However, because adjacent sections of the image overlap, processing each section individually results in a large amount of redundant computation and takes time.
[0011] Fully Convolutional Network (FCN) is another type of deep learning process. FCN can analyze an image and assign classification labels to each pixel within the image. As a result, compared with CNN, FCN is useful for analyzing images representing objects with two or more classifications. With FCN, an overlay map indicating the location of each classified object within the original image is generated. However, in order for the FCN deep learning algorithm to be effective, it is necessary to train on a dataset of images where each pixel is labeled as an organizational class, which takes too much annotation time and processing time to be practical. In digital WSI images, there may be cases where there are more than 10,000 to 100,000 pixels along each edge of the image. A complete image may contain at least 10,000 2 ~100,000 2 pixels, which requires a very long algorithm execution time to attempt to classify the tissue. Due to the large number of pixels, it is impossible to segment digital images of slides using conventional FCN.
[0012] There is a need for new technologies that can easily diagnose other biomarkers using TIL, PD-L1, and H&E images in order to efficiently identify and characterize such biomarkers across a population group, create more optimized drug treatment recommendations and protocols, and improve the prediction of disease progression. Summary of the Invention
[0013] This application presents an imaging-based biomarker prediction system formed by a deep learning framework configured and trained to directly learn from a pathological tissue slide image and predict the presence of biomarkers in a medical image. In an example, the deep learning framework is configured and trained to analyze a pathological tissue image and identify multiple different biomarkers. In various examples, these deep learning frameworks are configured to include different trained biomarker classifiers, and each of these biomarker classifiers is configured to receive an unlabeled pathological tissue image and provide different biomarker predictions for those images. These biomarker predictions can then be used to narrow a large set of available immunotherapies to a reduced small subset of targeted immunotherapies that can be used by medical professionals to treat patients. Thus, in various examples, a deep learning framework is provided that identifies a biomarker indicating the presence of a tumor, the state / condition of the tumor, or tumor-related information of a tissue sample, from which a set of targeted immunotherapies can be determined.
[0014] In an example, the system includes a deep learning framework that is trained to analyze and predict the biomarker status of a pathological tissue image received from a network-accessible image source, such as a medical laboratory or a medical imaging machine, and generate a report of the predicted biomarker status that can be saved and displayed. These predicted biomarker status reports are provided and saved and displayed in a network-accessible system, such as a pathology laboratory and a primary care physician system, and can be used to determine a patient's cancer treatment protocol (i.e., immunotherapy treatment or chemotherapy treatment). In some examples, the predicted biomarker status report can be input into a network-accessible next-generation sequencing system to drive subsequent genomic sequencing, or into a computerized cancer treatment decision system to filter a treatment list to corresponding treatments determined by the biomarker.
[0015] The technology of this specification can identify biomarkers associated with any of a wide variety of cancers. Exemplary cancers include, but are not limited to, adrenocortical carcinoma, lymphoma, anal cancer, anorectal cancer, basal cell carcinoma, cutaneous cancer (non-melanoma), cholangiocarcinoma, extrahepatic cholangiocarcinoma, intrahepatic cholangiocarcinoma, bladder cancer, urothelial bladder cancer, osteosarcoma, brain tumor, brainstem glioma, breast cancer (including triple-negative breast cancer), cervical cancer, colon cancer, colorectal cancer, lymphoma, endometrial cancer, esophageal cancer, gastric (stomach) cancer, head and neck cancer, hepatocellular (liver) cancer, kidney cancer, renal cell carcinoma, lung cancer, melanoma, tongue cancer, oral cancer, ovarian cancer, pancreatic cancer, prostate cancer, uterine cancer, testicular cancer, vaginal cancer.
[0016] In some examples, the imaging-based biomarker prediction system is formed by a deep learning framework, and the deep learning framework has a multi-scale configuration designed to perform the classification of (labeled or unlabeled) histopathological images using a classifier trained to classify tiles of the received histopathological images. In some examples, the multi-scale configuration includes a tile-level tissue classifier, i.e., a classifier trained using tile-based deep learning training. In some examples, the multi-scale configuration includes a pixel-level cell classifier and a cell segmentation model. In some examples, the classifications from the tile-level tissue classifier and the pixel-level cell classifier are analyzed to predict the status of biomarkers within the histopathological image. Further, in some examples, the multi-scale configuration includes a tile-level biomarker classifier.
[0017] In some examples, an imaging-based biomarker prediction system is formed by a deep learning framework, and the deep learning framework has a single-scale configuration designed to perform classification of (labeled or unlabeled) pathological tissue images using a classifier trained using multi-instance learning (MIL) techniques. In some examples, the single-scale configuration includes a slide-level classifier trained using gene sequencing data such as RNA sequencing data. That is, using RNA sequence data, a slide-level classifier is trained, and an image-based classifier capable of predicting the biomarker status of pathological tissue images is developed.
[0018] According to one example, a computer-implemented method for identifying biomarkers in a digital image of a hematoxylin and eosin (H&E) stained slide of a target tissue, the method comprising receiving the digital image at an image-based biomarker prediction system having one or more processors; performing an image tiling process on the digital image by using one or more processors to separate the digital image into a plurality of tile images, each of the plurality of tile images including a different portion of the digital image; applying the plurality of tile images to a multi-scale deep learning framework including one or more trained deep learning multi-scale classifier models, each trained to classify a different tissue classification for each tile image, and using the multi-scale deep learning framework to determine a tissue classification for each of the plurality of tile images; using one or more processors to identify cells in the digital image using a trained cell segmentation model; and identifying a predicted presence of one or more biomarkers associated with the digital image from the tissue classification determined for each tile image and from the identified cells in the digital image.
[0019] According to another example, a computer-implemented method for identifying biomarkers in digital images of hematoxylin and eosin (H&E) stained slides of a target tissue, the method comprising: receiving a molecular training data set of a plurality of training tissue samples, the molecular training data set including RNA transcriptome counts from sequencing of substantially similar samples associated with each training tissue sample; performing a clustering process on the molecular training data set to identify one or more molecular data subsets, each corresponding to a different biomarker; for each of the one or more molecular data subsets, receiving a plurality of digital images of H&E stained training slides of the training tissue samples corresponding to each biomarker for an image-based biomarker prediction system having one or more processors; using the one or more processors to generate a trained image-based biomarker classifier model for each of the one or more molecular data subsets based on the plurality of digital images of the H&E stained training slides; using the one or more processors to receive a subsequent digital image of an H&E stained slide of a subsequent tissue sample; and using the one or more processors to apply the subsequent digital image to the trained image-based biomarker classifier model to identify the predicted presence of one or more biomarkers in the subsequent tissue sample.
[0020] According to another example, a computer-implemented method for identifying biomarkers in a digital image of a hematoxylin and eosin (H&E) stained slide of a target tissue, the method comprising receiving the digital image at an image-based biomarker prediction system having one or more processors; using the one or more processors to separate the digital image into a plurality of tile images, each of the plurality of tile images including a different portion of the digital image; using the one or more processors to apply the plurality of tile images to a deep learning framework including one or more trained biomarker classification models, each trained to classify a different tissue classification; using the one or more processors to predict the biomarker classification of each of the plurality of tile images using the one or more trained biomarker classification models; determining a predicted presence of one or more biomarkers in the target tissue from the predicted biomarker classification of each of the tile images; and generating a report including the digital image and a digital overlay visualizing the predicted presence of the one or more biomarkers.
[0021] In some examples, the deep learning framework includes a multi-scale deep learning framework.
[0022] In some examples, separating the digital image into a plurality of tile images includes performing an image tiling process by applying a tiling mask to the digital image using the one or more processors to separate the digital image into a plurality of tile images.
[0023] In some examples, the tiling mask includes tiles of the same size and / or tiles having a rectangular shape.
[0024] In some examples, applying a plurality of tile images to a deep learning framework and predicting the biomarker classification of each of the plurality of tile images each involve applying each tile image to one or more trained deep learning multi-scale classifier models trained to classify different tissue classifications for each tile image, using a multi-scale deep learning framework to determine the tissue classification of each of the plurality of tile images, using one or more processors to identify cells in a digital image using a trained cell segmentation model, and predicting the biomarker classification of each tile image from the tissue classification determined for each tile image and from the identified cells in the digital image.
[0025] In some examples, the method further includes training one or more trained deep learning multi-scale classifier models in a multi-scale deep learning framework by receiving a plurality of H&E slide training images from a training image dataset, each H&E slide training image having a label corresponding to a biomarker to be trained, performing tile-based tissue classification analysis on each of the H&E slide training images, performing pixel-based cell segmentation analysis on each of the H&E slide training images, optionally performing tile-based biomarker classification analysis on each of the H&E slide training images, and generating one or more trained deep learning multi-scale classifier models accordingly.
[0026] In some examples, each H&E slide training image includes a plurality of tile images each having a tile-level label.
[0027] In some examples, the method includes, for each H&E slide training image, attaching a tile-level label to each of the plurality of tile images of the H&E slide training image.
[0028] In some examples, the method further includes, for each H&E slide training image, performing a tile selection process that infers the class status of each tile image within the H&E slide training image, and based on the inferred class status, discarding tile images that do not correspond to the target class before performing tile-based tissue classification analysis on each of the H&E slide training images, whereby the tile-based tissue classification analysis is performed only on the selected tile images of the H&E slide training images.
[0029] In some examples, one of the one or more trained deep learning multi-scale classifier models is configured as a fully convolutional network (FCN) classification model for tile resolution, respectively.
[0030] In some examples, identifying cells within a digital image tile using a trained cell segmentation model includes using one or more processors to apply each of a plurality of tile images to the cell segmentation model and, for each tile, assigning a cell classification to one or more pixels within the tile image.
[0031] In some examples, assigning a cell classification to one or more pixels within a tile image includes using one or more processors to identify one or more pixels as within a cell, at a cell boundary, or outside a cell, and classifying one or more pixels as within a cell, at a cell boundary, or outside a cell.
[0032] In some examples, the trained cell segmentation model is a three-dimensional UNet classification model at pixel resolution trained to classify within a cell, at a cell boundary, and outside a cell.
[0033] In some examples, one or more biomarkers are selected from the group consisting of tumor infiltrating lymphocytes (TIL), nuclear cytoplasmic ratio (NC), ploidy, signet ring morphology, and programmed death ligand 1 (PD-L1).
[0034] In some examples, the deep learning framework includes a single-scale deep learning framework.
[0035] In some examples, separating a digital image into a plurality of tile images includes performing an image tiling process by applying the digital image to a trained multi-instance learning controller that separates the digital image into a plurality of tile images using one or more processors.
[0036] In some examples, the method further includes providing each tile image to a tile selection process that infers the class status of each tile image in the H&E slide training image, and selectively discarding tile images based on tile selection criteria before applying the remaining plurality of tile images to a deep learning framework based on the inferred class status.
[0037] In some examples, the method further includes providing each tile image to a tile selection process that infers the class status of each tile image in the H&E slide training image, and randomly discarding tile images before applying the remaining plurality of tile images to a deep learning framework.
[0038] In some examples, the method comprises receiving a molecular training dataset of a plurality of training tissue samples, wherein the molecular training dataset comprises RNA transcriptome counts from sequencing of substantially similar samples associated with each training tissue sample, performing a clustering process on the molecular training dataset to identify one or more molecular data subsets, each corresponding to a different biomarker, for each of the one or more molecular data subsets, receiving a plurality of digital images of H&E stained training slides of the training tissue samples corresponding to each biomarker for an image-based biomarker prediction system having one or more processors, and using the one or more processors to generate one of the trained biomarker classification models for each of the one or more molecular data subsets based on the plurality of digital images of the H&E stained training slides.
[0039] In some examples, for each of the one or more molecular data subsets, generating one of the trained biomarker classification models comprises performing a multi-instance learning process on the plurality of digital images of the H&E stained training slides.
[0040] In some examples, each of the plurality of digital images of the H&E stained training slides of the training tissue samples has a slide-level label.
[0041] In some examples, each of the plurality of digital images of the H&E stained training slides of the training tissue samples is unlabeled.
[0042] In some examples, the single-scale deep learning framework is a convolutional neural network having a ResNet configuration or an Inception-v3 configuration.
[0043] In some examples, one or more biomarkers are selected from the group consisting of consensus molecular subtype (CMS) and homologous recombination deficiency (「HRD」).
[0044] In some examples, one or more processors are one or more graphics processing units (GPUs), tensor processing units (TPUs), and / or central processing units (CPUs).
[0045] In some examples, a computing device (e.g., an image-based biomarker prediction system) is communicatively coupled to a pathology slide scanner system via a communication network, whereby the image-based biomarker prediction system receives a digital image from the pathology slide scanner system via the communication network.
[0046] In some examples, the computing device is included within the pathology slide scanner system.
[0047] In some examples, the pathology slide scanner system includes an image-based, adversarially trained, and / or microsatellite instability (MSI) prediction model.
[0048] In some examples, generating a report that includes a digital image and a digital overlay includes generating the digital overlay to include an overlay element that identifies the tumor content or tumor percentage of the digital image.
[0049] According to another example, a computing device configured to identify biomarkers in a digital image of a hematoxylin and eosin (H&E) stained slide of a target tissue, the computing device comprising: one or more memories; and one or more processors, the one or more processors being configured to: receive the digital image; perform an image tiling process on the digital image by separating the digital image into a plurality of tile images, each of the plurality of tile images including a different portion of the digital image; apply the plurality of tile images to a multi-scale deep learning framework including one or more trained deep learning multi-scale classifier models each trained to classify a different tissue classification for each tile image, and use the multi-scale deep learning framework to determine the tissue classification of each of the plurality of tile images; identify cells in the digital image using a trained cell segmentation model; and identify the predicted presence of one or more biomarkers associated with the digital image from the tissue classification determined for each tile image and from the identified cells in the digital image.
[0050] According to another example, a computing device configured to identify biomarkers in a digital image of a hematoxylin and eosin (H&E) stained slide of a target tissue, the computing device comprising: one or more memories; and one or more processors configured to: receive a molecular training dataset of a plurality of training tissue samples, the molecular training dataset including RNA transcriptome counts from sequencing of substantially similar samples associated with each training tissue sample; perform a clustering process on the molecular training dataset to identify one or more molecular data subsets, each corresponding to a different biomarker; for each of the one or more molecular data subsets, receive a plurality of digital images of H&E stained training slides of the training tissue samples corresponding to each biomarker for an image-based biomarker prediction system having one or more processors; for each of the one or more molecular data subsets, generate a trained image-based biomarker classifier model based on the plurality of digital images of the H&E stained training slides; receive a subsequent digital image of an H&E stained slide of a subsequent tissue sample; and apply the subsequent digital image to the trained image-based biomarker classifier model to identify the predicted presence of one or more biomarkers in the subsequent tissue sample.
[0051] According to another example, a computing device configured to identify biomarkers in a digital image of a hematoxylin and eosin (H&E) stained slide of a target tissue, the computing device comprising: one or more memories; and one or more processors, the one or more processors being configured to: receive the digital image into an image-based biomarker prediction system; separate the digital image into a plurality of tile images, each of the plurality of tile images including a different portion of the digital image; apply the plurality of tile images to a deep learning framework including one or more trained biomarker classification models, each of the one or more trained biomarker classification models being trained to classify a different tissue classification; predict the biomarker classification of each of the plurality of tile images using the one or more trained biomarker classification models; determine a predicted presence of one or more biomarkers in the target tissue from the predicted biomarker classifications of each of the tile images; and generate a report including the digital image and a digital overlay visualizing the predicted presence of the one or more biomarkers.
Brief Description of the Drawings
[0052] This patent or the patent application file contains at least one drawing executed in color. Copies of this patent or patent application publication with color drawings will be provided by the United States Patent and Trademark Office upon request and payment of the necessary fee.
[0053] The drawings described below illustrate various aspects of the systems and methods disclosed herein. It should be understood that each drawing illustrates an example of an aspect of the systems and methods.
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10A
Figure 10B
Figure 11
Figure 12A
Figure 12B
Figure 12C
Figure 13
Figure 14
Figure 15A
Figure 15B
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21
Figure 22
Figure 23
Figure 24
Figure 25
Figure 26
Figure 27
Figure 28
Figure 29
Figure 30
Figure 31
Figure 32
Figure 33
Figure 34
Figure 35
Figure 36
Figure 37
Figure 38
DETAILED DESCRIPTION OF THE INVENTION
[0054] The imaging-based biomarker prediction system is composed of a deep learning framework configured and trained to directly learn from pathological tissue slides and predict the presence of biomarkers in medical images. The deep learning framework can be configured and trained to analyze medical images and identify biomarkers indicating the presence of tumors, the state / condition of tumors, or information about tumors in tissue samples.
[0055] In an implementation form, a cloud-based deep learning framework is used for medical image analysis. The deep learning algorithm automatically learns advanced imaging functions to enhance diagnosis, prognosis, treatment indication, and prediction of treatment effects. In an example, the deep learning framework can be directly connected to cloud storage and utilize resources on the cloud platform to perform efficient training, comparison, and deployment of deep learning algorithms.
[0056] In some examples, the deep learning framework includes a multi-scale configuration that uses a tiling strategy to accurately capture the structural and local tissue structures of various diseases (e.g., predict cancer tumors). These multi-scale configurations perform classification of pathological tissue images (with or without labels) using a classifier trained to classify tiles of the received pathological tissue images. In some examples, the multi-scale configuration includes a tile-level tissue classifier, i.e., a classifier trained using tile-based deep learning training. In some examples, the multi-scale configuration includes a pixel-level cell classifier and a cell segmentation model. In some examples, the classification from the tile-level tissue classifier and the pixel-level cell classifier is analyzed to predict the status of biomarkers in the pathological tissue image. Further, in some examples, the multi-scale configuration includes a tile-level biomarker classifier. When tracked, the multi-scale classifier can receive new labeled or unlabeled pathological tissue images and predict the presence of specific biomarkers in the associated pathological tissue slides.
[0057] In some examples, the deep learning framework of this specification includes a single-scale configuration trained using a multi-instance learning (MIL) strategy to predict the presence of biomarkers in pathological tissue images. A classifier trained using a single-scale configuration can be trained to perform classification of (labeled or unlabeled) pathological tissue images using a classifier trained using one or more multi-instance learning (MIL) techniques. In some examples, the single-scale configuration includes a slide-level classifier, and the slide-level classifier is trained using gene sequencing data such as RNA sequencing data and is trained to analyze pathological tissue images having slide-level labels rather than tile-level labels. That is, using RNA sequence data, a slide-level classifier is trained, and an image-based classifier capable of predicting the biomarker status of pathological tissue images is developed.
[0058] Any of the multi-scale and single-scale configurations of this specification can incorporate optimizations of various algorithms to accelerate the calculations for such disease analysis.
[0059] In an implementation of the multi-scale classifier configuration, the deep learning framework can be trained to include a classifier that performs automatic cell segmentation, determines the cell / biomarker type, determines the tissue type classification from the pathological tissue image, thereby providing image-based biomarker development. Even a single-scale classifier can be trained to include tissue type classification and biomarker classification.
[0060] In the case of a multi-scale classifier configuration, for example, the intensive and spatial imaging functions for various cell types (e.g., tumor, stroma, lymphocyte) of digital hematoxylin & eosin (H&E) slides are determined by a deep learning framework and can be used to predict clinical and treatment outcomes. Instead of a preliminary manual cell type classification, in the examples herein, a deep learning framework uses a multi-scale configuration to classify each sub-region of an H&E slide histopathological image into specific cell segmentation, cell type, and tissue type. From there, biomarker detection is performed by another deep learning framework configured to identify various types of imaging metrics. Examples of imaging metrics include tumor shape, including the minimum and maximum shape of the tumor, tumor area, tumor perimeter, tumor %, cell shape, which includes cell area, cell perimeter, cell convex area ratio, cell circularity, cell convex perimeter area, cell length, lymphocyte %, cell characteristics, cell texture, and cell texture includes saturation, intensity, and hue.
[0061] Examples of tissue classes include, but are not limited to, tumor, stroma, normal, lymphocyte, fat, muscle, blood vessel, immune cluster, necrosis, hyperplasia / dysplasia, erythrocyte, and tissue classes or cell types that are positive (specifically, including amounts of the IHC staining target molecule greater than a specific threshold) or negative (not including the molecule or including amounts of the molecule less than a specific threshold) for an IHC staining target molecule.
[0062] In some examples, biomarker detection can be enhanced by combining imaging metrics with structured clinical and sequencing data to develop enhanced biomarkers.
[0063] Biomarkers can be identified by any of the following models. Any model referred to herein can be implemented as an artificial intelligence engine and may include a gradient boosting model, a random forest model, a neural network (NN), a regression model, a naive Bayes model, or a machine learning algorithm (MLA). The MLA or NN can be trained with a training dataset. In an exemplary prediction profile, the training dataset may include imaging, medical conditions, clinical, and / or molecular reports, as well as patient details (e.g., curated from an EHR or a gene sequencing report). The MLA includes linear regression, logistic regression, decision trees, classification and regression trees, naive Bayes, supervised algorithms using nearest neighbor clustering (such as algorithms annotated with features / categories in the dataset), and unsupervised algorithms using Apriori, mean clustering, principal component analysis, random forest, adaptive boosting (such as algorithms not annotated with features / categories in the dataset), generative approaches (such as mixtures of Gaussian distributions, mixtures of multinomial distributions, hidden Markov models), low density separation, graph-based approaches (such as minimum cut, harmonic functions, manifold regularization), heuristic approaches, or semi-supervised algorithms using support vector machines (such as algorithms annotated with an incomplete number of features / categories in the dataset). The NN includes conditional random fields, convolutional neural networks, attention-based neural networks, deep learning, long short-term memory networks, or other neural models, and the training dataset includes a plurality of tumor samples, RNA expression data for each sample, and a pathology report covering imaging data for each sample. The MLA and neural networks identify different approaches to machine learning, but these terms can be used interchangeably herein. Thus, unless otherwise specified, a reference to the MLA may include the corresponding NN, and a reference to the NN may include the corresponding MLA.Training may include providing an optimized dataset, labeling these characteristics seen in patient records, and training the MLA to predict or classify based on new inputs. Artificial NNs are efficient computational models and are strong in solving difficult problems of artificial intelligence. It has been shown that artificial NNs are universal approximators (able to represent a wide variety of functions when given appropriate parameters). Some MLAs can identify important features and their coefficients or weights. The score can be generated by multiplying the coefficient by the frequency of occurrence of the feature, and when the score of one or more features exceeds a threshold, a specific classification can be predicted by the MLA. The coefficient schema can be combined with a rule-based schema to generate more complex predictions, such as predictions based on multiple features. For example, 10 main features may be identified in various classifications. There may be a list of coefficients for the main features and a set of rules for classification. The set of rules may be based on the occurrence count of the feature, the scaled weight of the feature, or other qualitative and quantitative evaluations of the feature encoded in logic known to those skilled in the art. In other MLAs, the features may be organized in a binary tree structure. For example, the main feature that can distinguish most classifications may exist as the root of the binary tree and each subsequent branch in the tree until a classification is assigned based on reaching the terminal node of the tree. For example, the binary tree may have a root node that tests the first feature. The occurrence or non-occurrence of this feature must exist (binary decision), and the logic can traverse the branch that is true for the item being classified. Additional rules may be based on thresholds, ranges, or other qualitative and quantitative tests. The supervised method is useful when there are many known values or annotations in the training dataset, but due to the nature of EMR / EHR documents, there may not be many provided annotations. When exploring large amounts of unlabeled data, the unsupervised method is useful for binning / bucketing the instances in the dataset.In this specification, a single instance of the above model, or a combination of two or more such instances, can be used to construct a model for the purpose of a model, artificial intelligence, neural network, or machine learning algorithm.
[0064] In some examples, the present technology provides machine learning-assisted pathological tissue image review, which includes automatically identifying and contouring tumor regions and / or characteristics of regions or cell types within the regions (e.g., lymphocytes, PD-L1 positive cells, tumors with high tumor budding, etc.), counting the cells within the tumor regions, and generating a decision score to improve the efficiency and objectivity of pathological slide review.
[0065] As used herein, the term "biomarker" refers to information derived from images related to the screening, diagnosis, prognosis, treatment, selection, disease monitoring, progression, and disease recurrence of cancer or other diseases, particularly information in the form of morphological features identifiable in histologically stained samples. In some examples, the biomarkers herein can be those of morphological features determined from labeled-based images. The biomarkers herein can be those of morphological features determined from labeled RNA data.
[0066] The biomarkers herein can be information derived from images that correlate with the presence of cancer or susceptibility to cancer in a subject, the likelihood that the cancer is one subtype or another, the presence or proportion of biological characteristics such as tissue, cell, or protein types or classes, the probability that a patient will respond or not respond to a particular treatment or class of treatments, the degree of positive response expected from a treatment or class of treatments (e.g., survival period and / or progression-free survival period), whether a patient is responding to treatment, or the likelihood that the cancer will regress, progress, or progress beyond the site of origin (i.e., metastasize).
[0067] Examples of biomarkers predicted from pathological tissue images using the various techniques of this specification include the following.
[0068] As used herein, tumor infiltrating lymphocytes (TILs) refer to mononuclear immune cells that infiltrate tumor tissue or stroma. TILs include, for example, T cells, B cells, and NK cells, and these populations can be further classified based on function, activity, and / or biomarker expression. For example, populations of TILs can include, for example, cytotoxic T cells that express CD3 and / or CD8, and regulatory T cells (also known as suppressor T cells), which are often characterized by FOXP3 expression. Information regarding the density, location, tissue, and composition of TILs provides valuable insights into prognosis and potential treatment options. The present disclosure provides, in various aspects, methods for predicting TIL density in a sample, methods for distinguishing subpopulations of TILs in a sample (e.g., distinguishing cytotoxic T cells expressing CD3 / CD8 from FOXP3 Tregs), methods for distinguishing stroma from intratumoral TILs, and the like.
[0069] Programmed death ligand 1 (PD-L1) is a 40 kDa type I transmembrane protein that affects immune system suppression, particularly in patients with autoimmune diseases, cancer, and other pathologies. In the context of cancer immunotherapy, PD-L1 is expressed on the surface of tumor cells, tumor-associated macrophages (TAMs), and T lymphocytes, and can subsequently inhibit PD-1 positive T cells.
[0070] Ploidy refers to the number of sets of homologous chromosomes within the genome of a cell or organism. Examples include haploid, which means one set of chromosomes, and diploid, which means two sets of chromosomes. The presence of multiple sets of chromosomes that pair with the genome of an organism is described as "polyploid." Three sets of chromosomes, 3n, is triploid, and four sets of chromosomes, 4n, is tetraploid. Extremely large numbers of sets can be designated by number (e.g., 15 sets would be 15-ploid).
[0071] The nuclear-to-cytoplasmic (NC) ratio is a measure of the ratio of the size of a cell's nucleus to the size of its cytoplasm. The NC ratio can be expressed as a volume ratio or a cross-sectional area. The NC ratio can indicate the maturity of a cell, and the size of the cell nucleus decreases with cell maturity. In contrast, a high NC ratio in a cell may indicate a malignant tumor of the cell.
[0072] The signet-ring morphology is the morphology of signet-ring cells, i.e., cells with large vacuoles, and its malignant form mainly appears in the case of carcinomas. Signet-ring cells are most highly associated with gastric cancer, but can arise from various tissues such as the prostate, bladder, gallbladder, breast, colon, ovarian stroma, testis, etc. For example, signet-ring cell carcinoma (SRCC) is a rare form of highly malignant adenocarcinoma. This is an epithelial malignant tumor characterized by the histological appearance of signet-ring cells.
[0073] These biomarkers, TIL, NC ratio, ploidy, signet-ring morphology, and PD-L1 are examples of biomarkers of morphological features determined from labeled-based images by the techniques of this specification.
[0074] Consensus molecular subtypes ("CMS") is a set of classification subtypes of colorectal cancer (CRC) developed based on comprehensive gene expression profile analysis. CMS classification in primary colorectal cancer includes CMS1 - immune infiltration (often BRAFmut, MSI-High, TMB-High), CMS2 - canonical (often ERBB / MYC / WNT-driven), CMS3 - metabolic (often KRASmut), and CMS4 - mesenchymal (often TGF-B-driven). More broadly, the CMS of this specification includes these and other subtypes of colorectal cancer. Even more broadly, the CMS of this specification refers to subtypes derived from comprehensive gene expression profile analysis of other cancer types described in this specification.
[0075] The same homologous recombination deficiency (「HRD」) status is a classification indicating a defect in the normal homologous recombination DNA damage repair process that results in the loss of duplication of chromosomal regions referred to as heterozygous genomic loss (LOH).
[0076] Biomarkers such as CMS and HRD are examples of biomarkers of morphological characteristics determined from labeled RNA data by the techniques herein.
[0077] By way of example, the biomarkers herein include HRD status, DNA ploidy score, karyotype, CMS score, chromosomal instability (CIN) status, ring form score, NC ratio, activation status of cellular pathways, cell state, tumor characteristics, and splice variants.
[0078] As used herein, 「pathological tissue image」 refers microscopically to a digital (including digitized) image of histopathologically developed tissue. Examples include images of histologically stained specimen tissues, where histological staining is a process performed in the preparation of sample tissues to assist microscopic examination. In some examples, the pathological tissue image is a digital image of a hematoxylin and eosin (H&E) stained pathological tissue slide, an immunohistochemical (IHC) stained slide, a Romanowsky - Giemsa stained slide, a Gram stained slide, a trichrome stained slide, a carmine stained slide, and a silver nitrate stained slide. Other examples include blood smear slides and tumor smear slides. In other examples, the pathological tissue image is that of other stained slides known in the art. As used herein, references to digital images, digitized images, slide images, and medical images refer to 「pathological tissue image」.
[0079] These histopathological images can be captured in the visible wavelength region and beyond, such as infrared digital images obtained using spectroscopy of histopathologically developed tissues. In some examples, the histopathological images include z-stack images representing horizontal cross-sections of three-dimensional specimens or histopathological slides captured at various levels of the specimen or at various foci of the slide. In some examples, two or more images can be images from adjacent or nearly adjacent sections of tissue from the specimen, and one of the two or more images can have a tissue feature corresponding to a different tissue feature of the two or more images. There can be a vertical and / or horizontal shift between the position of the corresponding tissue feature in the first image and the position of the corresponding tissue feature in the second image. Thus, histopathological images also refer to images, sets of images, or videos generated from multiple different images. It should be understood that the following exemplary embodiments can be stained, exchanged, or trained with models in different styles of staining, unless explicitly excluded.
[0080] Various examples in this specification are described with reference to a particular class of histopathological images, H&E slide images. Digital H&E slide images can be generated by capturing digital photographs of H&E slides. Alternatively, or in addition, such images can be generated from images derived from unstained tissue via a machine learning system such as deep learning. For example, a digital H&E slide image may be generated from a wide-field autofluorescence image of an unlabeled tissue section. See, for example, Rivenson et al., "Virtual histological staining of unlabelled tissue-autofluorescence images via deep learning", Nature Biomedical Engineering, 3(6):466, 2019.
[0081] FIG. 1 shows a prediction system 100 that can analyze a digital image of a pathological tissue slide of a tissue sample and determine the likelihood of the presence of a biomarker in the tissue, where the presence of the biomarker indicates other information about the tumor in the tissue sample, such as the presence of a predicted tumor, the predicted state / condition of the tumor, or the likelihood of a clinical response by the use of a treatment related to the biomarker.
[0082] System 100 includes an imaging-based biomarker prediction system 102 that implements, in particular, image processing operations, a deep learning framework, and a report generation operation to analyze the pathological tissue image of the tissue sample and predict the presence of a biomarker in the tissue sample. In various examples, system 100 is configured to predict the presence of these biomarkers, the location of the tissue associated with these biomarkers, and / or the location of the cells of these biomarkers.
[0083] The imaging-based biomarker prediction system 102 can be implemented on one or more computing devices, such as a computer, a tablet, or other mobile computing device, or on a server such as a cloud server. The imaging-based biomarker prediction system 102 can include, as described herein, a plurality of processors, controllers, or other electronic components for processing or facilitating the capture, generation, or storage of images and image analysis, as well as deep learning tools for the analysis of images. An exemplary computing device 3800 for implementing the imaging-based biomarker prediction system 102 is shown in FIG. 38.
[0084] As shown in FIG. 1, an imaging-based biomarker prediction system 102 is connected to one or more medical data sources via a network 104. The network 104 can be a public network such as the Internet, a private network such as a research institution or corporate private network, or any combination thereof. The network includes local area networks (LANs), wide area networks (WANs), cellular, satellite, or other network infrastructure, whether wireless or wired. The network 104 can be part of a cloud-based platform. The network 104 can utilize packet-based and / or datagram-based communication protocols such as Internet Protocol (IP), Transmission Control Protocol (TCP), User Datagram Protocol (UDP), or other types of protocols. Additionally, the network 104 can include a plurality of devices that facilitate network communication and / or form the hardware infrastructure of the network, such as switches, routers, gateways, access points (such as wireless access points as shown), firewalls, base stations, repeaters, backbone devices, etc.
[0085] The imaging-based biomarker prediction system 102 is communicatively coupled via a network 104 to receive pathological tissue slides such as medical images, e.g., digital H&E stained slide images, IHC stained slide images, or digital images of other staining protocols from a wide variety of different sources. These sources can include a physician's clinical record system 106 and a pathological tissue imaging system 108. The system 100 can be used to access any number of medical image data sources. The pathological tissue images may be images captured by any suitable optical pathological tissue slide scanner, including any dedicated digital medical image scanner, e.g., a magnification scanner with resolutions of 20x and 40x. Further, the biomarker prediction system 102 can receive images from a pathological tissue image repository 110. In yet other examples, the images may be received from a partner genomic sequencing system 112, e.g., TCGA and NCI Genomic Data Commons. Further, the biomarker prediction system 102 can receive pathological tissue images from an organoid modeling lab 116. These image sources can communicate image data, genomic data, patient data, treatment data, historical data, etc., according to the techniques and processes described herein. Each of the image sources can represent a plurality of image sources. Further, each of these image sources can be considered a different data source, and those data sources may be capable of generating and providing different image data than other providers, hospitals, etc. The imaging data between different sources can potentially differ in one or more respects, resulting in various data source-specific biases such as different dyes, fixation of biological samples, embedding, staining protocols, as well as different pathological imaging devices and settings.
[0086] In the example of FIG. 1, the imaging-based biomarker prediction system 102 includes an image preprocessing subsystem 114, which performs initial image processing to enhance the image data in order to speed up the processing in the training of a machine learning framework and perform biomarker prediction using a trained deep learning framework. In the illustrated example, the image preprocessing subsystem 114 performs a normalization process including one or more of color normalization 114a, intensity normalization 114b, and imaging source normalization 114c on the received image data to compensate for and correct differences in the received image data. In some examples, the imaging-based biomarker prediction system 102 receives medical images, while in other examples, the subsystem 114 can generate medical images from either the received pathological tissue slides or other received images, for example, by aligning shifted pathological tissue images to compensate for vertical / horizontal shifts to generate composite pathological tissue images. This preprocessing of the images enables the deep learning framework to more efficiently analyze images over large datasets (e.g., more than 1000, more than 10000, more than 100000, more than 1000000 medical images), thus speeding up the training and analysis processes.
[0087] The image preprocessing subsystem 114 performs further image processing to remove artifacts and other noise from the received images by performing preliminary tissue detection 114d, and can identify regions of the image corresponding to histopathologically stained tissue, for example, for subsequent analysis, classification, and segmentation.
[0088] As further described herein, in a multi-scale configuration where the image data is analyzed on a tile basis, in some examples, the preprocessing of the images includes receiving an initial pathological tissue image at a first image resolution, downsampling the image to a second image resolution, and then performing normalization on the downsampled pathological tissue image, such as color and / or intensity normalization, to remove non-tissue objects from the image.
[0089] In contrast, in the single-scale configuration, downsampling of the received pathological tissue image is not used. The single-scale configuration analyzes image data at the slide level rather than at the tile level.
[0090] In still further hybrid versions of each of the multi-scale and single-scale configurations, a tiling process is applied to the received pathological tissue image to generate tiles for tile-based analysis.
[0091] The imaging-based biomarker prediction system 102 can be a stand-alone system that interfaces with external (i.e., third-party) network-accessible systems 106, 108, 110, 112, and 116. In some examples, the imaging-based biomarker prediction system 102 can be integrated with one or more of these systems as part of a distributed cloud-based platform. For example, system 102 can be integrated with a pathological tissue imaging system such as a digital H&E staining imaging system, enabling, for example, rapid biomarker analysis and reporting at the imaging station. Indeed, any of the functions described herein using the techniques can be distributed across one or more network-accessible devices, including cloud-based devices.
[0092] In some examples, the imaging-based biomarker prediction system 102 is part of an encompassing biomarker prediction, patient diagnosis, and patient treatment system. For example, the imaging-based biomarker prediction system 102 can be coupled to communicate predicted biomarker information, tumor predictions, and tumor status information to an external system, which includes a computer-based pathology laboratory / oncology system 118 that can receive a generated biomarker report including image overlay mapping and use it to further diagnose a patient's cancer condition and identify corresponding treatments for use in the patient's treatment. The imaging-based biomarker prediction system 102 can further transmit the generated report to the patient's primary care provider's computer system 120 and the physician's clinical record system 122 for databaseing the patient report using a database of past generated reports regarding the patient and / or generated reports regarding other patients for use in future patient analysis (including the deep learning analysis described herein).
[0093] To analyze the received histopathological image data and other data, the imaging-based biomarker prediction system 102 includes a deep learning framework 150 that implements various machine learning techniques to generate a trained classifier model for image-based biomarker analysis from a set of received image data or a set of image data and other patient information. Using the trained classifier model, the deep learning framework 150 is further used to analyze and diagnose the presence of image-based biomarkers in subsequent images collected from a patient. In this way, images and other data of patients previously treated and analyzed are utilized through the trained model to provide analysis and diagnostic capabilities for future patients.
[0094] In an exemplary system 100, the deep learning framework 150 includes a pathological tissue image-based classifier training module 160, which can access data received and stored from external systems 106, 108, 110, 112, and 116, as well as any other systems, where the data is analyzed from the received data stream and can be database-ized into various data types. These various data types can be divided into image data 162a, which can be associated with other data types such as molecular data 162b, demographic data 162c, and tumor response data 162d. The association can be formed by labeling the image data 162a with one or more different data types. By labeling the image data 162a according to the association with other data types, the imaging-based biomarker prediction system can train the image classifier module to predict one or more different data types from the image data 162a.
[0095] In the illustrated data, the deep learning framework 150 includes the image data 162a. For example, to train or use a multi-scale PD-L1 biomarker classifier, this image data 162a can include pre-processed image data received from the subsystem 114, images from H&E slides, or images from IHC slides (with or without human annotation) targeting hormone receptors such as PD-L1, PTEN, EGFR, beta-catenin / catenin beta 1, NTRK, HRD, PIK3CA, and HER2, AR, ER, PR. Whether it is a multi-scale classifier or a single-scale classifier, for training or using other biomarker classifiers, the image data 162A may include images from other stained slides. Further, in an example of training a single-scale classifier, the image data 162A is image data related to RNA sequence data of a specific biomarker cluster that enables the multi-instance learning (MIL) technique herein.
[0096] Molecular data 162b can include DNA sequences, RNA sequences, metabolomics data, proteomics / cytokine data, epigenomic data, organoid data, karyotype data, transcriptional data, transcriptomics, metabolomics, microbiomics, and immunomics, and can include the identification of SNPs, MNPs, InDels, MSIs, TMBs, CNV fusions, loss of heterozygosity, loss or gain of function. Epigenomic data includes DNA methylation, histone modification, or other factors that inactivate genes without changing the nucleotide sequence of the gene or cause changes in gene function. Microbiomics includes data on viral infections that can affect the treatment and diagnosis of specific diseases, as well as data on bacteria present in a patient's gastrointestinal tract that can affect the effectiveness of drugs the patient takes. Proteomics data includes the composition, structure, and activity of proteins, when and where proteins are expressed, the rates of protein production, degradation, and steady-state abundance, how proteins are modified (e.g., post-translational modifications such as phosphorylation), the movement of proteins between intracellular compartments, the involvement of proteins in metabolic pathways, how proteins interact with each other, or post-translational modifications of proteins from RNA, such as phosphorylation, ubiquitination, methylation, acetylation, glycosylation, oxidation, or nitrosylation.
[0097] Deep learning framework 150 can further include demographic data 162c and tumor response data 162d (e.g., data regarding the reduction in tumor growth after being subjected to immunotherapy, DNA damage therapies such as PARP inhibitors or platinum, or specific therapies such as HDAC inhibitors). Demographic data 162c can include age, gender, race, country of origin, etc. Tumor response data 162d can include epigenomic data, examples of which include changes in chromatin conformation and histone modification.
[0098] Tumor response data 162d may include cellular pathways, examples of which include the IFN gamma, EGFR, MAP kinase, mTOR, CYP, CIMP, and AKT pathways, as well as pathways downstream of HER2 and other hormone receptors. Tumor response data 162d may include cellular state indicators, examples of which include the composition, appearance, or refractive index of collagen (e.g., extracellular vs. fibroblasts, nodular fasciitis), the density of the stroma or other stromal characteristics (e.g., stromal thickness, wet vs. dry), and / or angiogenesis or the general appearance of the vasculature (including the distribution of the vasculature in the collagen / stroma, also referred to as epithelial-mesenchymal transition or EMT). Tumor response data 162d may include tumor characteristics, examples of which include the complexity of the tumor, the presence of tumor budding or other morphological features / characteristics indicating the size of the tumor (including the bulky or lean state of the tumor), the aggressiveness of the tumor (e.g., high-grade basaloid tumors, particularly in colorectal cancer, or high-grade dysplasia, particularly known in Barrett's esophagus), and / or the immune state of the tumor (e.g., contrast between inflammatory / "hot" tumors and non-inflammatory / "cold" tumors and immune-excluded tumors).
[0099] The histopathology image-based classifier training module 160 may be composed of machine learning techniques adapted for image analysis, including, for example, deep learning techniques, which include, as an example, a CNN model, more specifically, a CNN at tile resolution, which in some examples is implemented as an FCN model, more specifically, an FCN model at tile resolution. Any of the data types 162a - 162d are included within the histopathology image and can be directly obtained from data transmitted to the imaging-based biomarker prediction system 102, such as data transmitted with the histopathology image. The data types 162a - 162d may be used by the histopathology image-based classifier training module 160 to develop a classifier for identifying one or more biomarkers discussed herein.
[0100] In one example, a pathological tissue image can be segmented, and each segment of the image can be labeled according to one or more data types into which the segment can be classified. In another example, a pathological tissue image can be labeled as a whole according to one or more data types into which the image or at least one segment of the image can be classified. The data type can indicate one or more biomarkers, and by labeling the pathological tissue image or segment with the data type, the biomarker can be identified.
[0101] In an exemplary system 100, the deep learning framework 150 can further include a trained image classifier module 170 composed of deep learning techniques, which includes those that implement the module 160. In some examples, the trained image classifier module 170 accesses the image data 162 for analysis and biomarker classification. In some examples, the module 170 further accesses the molecular data 162, demographic data 162c, and / or tumor response data 162d for analysis and tumor prediction, treatment response prediction, etc.
[0102] The trained image classifier module 170 includes a trained tissue classifier 172, and the trained tissue classifier 172 is trained by the module 160 using one or more training image sets to identify and classify tissue types in the received image data region / area. In some examples, these trained tissue classifiers are trained to identify biomarkers through tissue classification, and these classifiers include a single-scale component classifier 172a and a multi-scale classifier 172b.
[0103] The module 170 can further include other trained classifiers, which include a trained cell classifier 174 that identifies biomarkers through cell classification. The module 170 can further include a cell segmenter 176 that identifies cells within a pathological tissue image, including cell boundaries, the interior of the cells, and the exterior of the cells.
[0104] In the examples of this specification, the tissue classifier 172 may include a biomarker classifier specially trained to identify, according to the biomarkers of this specification, tumor infiltration (such as the ratio of lymphocytes in the tumor tissue to all cells in the tumor tissue), PD-L1 (such as positive or negative status), ploidy (such as score), CMS (such as identification of subtypes), NC ratio (such as identification of nuclear size), signet ring morphology (such as classification of signet cells and vacuole size), HRD (such as by score, or positive or negative classification), and the like.
[0105] As detailed herein, the trained image classifier module 170 and related classifiers may be composed of machine learning techniques adapted for image analysis, including, for example, deep learning techniques, and the deep learning techniques include, by way of example, a CNN model, more specifically, a CNN at tile resolution, which in some examples is implemented as an FCN model, and more specifically, implemented as an FCN model at tile resolution, and the like.
[0106] The system 102 further includes a tumor report generator 180 configured to receive classification data from the trained tissue (biomarker) classifier 172, the trained cell (biomarker) classifier 174, and the cell segmenter 172, determine tumor metrics of the image data, and generate a digital image and a statistical data report. Here, such output data may be provided to the pathology laboratory 118, the primary care physician system 120, the genomic sequencing system 112, the tumor board, the tumor board electronic software system, or other external computer systems for display or consumption in further processes.
[0107] A conventional cancer diagnosis workflow 200 using pathological tissue images is shown in FIG. 2. A biopsy is performed to collect tissue samples from a patient. In a medical laboratory, digital pathological tissue images of the tissue samples are generated (202) using known staining techniques such as H&E or IHC staining and a digital medical imager (e.g., a slide scanner). These pathological tissue images are provided to a pathologist who visually analyzes them, and tumors within the received images are identified (204). The pathologist can optionally receive the patient's genomic sequencing data (e.g., DNA Seq data or RNA Seq data from a genomic sequencing lab) and analyze that data (206). Next, the pathologist diagnoses other characteristics of the type of cancer of the tumor / cancer cells from the visual analysis of the pathological tissue slides and any genomic sequencing data (208) and creates a pathology report (210).
[0108] FIG. 3 shows an exemplary implementation of a deep learning framework 150 in the form of an imaging-based biomarker prediction system 102, more specifically, a deep learning framework 300. The framework 300 can be communicatively coupled to receive pathological tissue image data and other data (molecular data, tumor response data, demographic data, etc.) via a network 104 from external systems such as a physician's clinical record system 106, a pathological tissue imaging system 108, a genomic sequencing system 112, a medical image repository 110, and / or the organoid modeling lab 116 of FIG. 1. The organoid modeling lab 116 can collect various types of data, such as the sensitivity of the organoids to a drug (determined, for example, by measuring cell death or cell viability after exposure to the drug), single-cell analysis data, or the detection of cell products (including proteins, lipids, and other molecules) indicating the presence of a particular cell population, which can include effector data, stimulus data, regulatory data, inflammatory data, chemotactic data, as well as organoid image data, any of which can be stored within the molecular data 162b.
[0109] The framework 300 includes a preprocessing controller 302, a deep learning framework cell segmentation module 304, a deep learning framework multi-scale classifier module 306, a deep learning framework single-scale classifier module 307, and a deep learning postprocessing controller 308.
[0110] To prepare medical images for multi-scale and single-scale deep learning, in one example, the preprocessing controller 302 includes a normalization process 310, which may include color normalization, intensity normalization, and imaging source normalization. The normalization process 310 is optional and can be excluded to facilitate deep learning training, image analysis, and / or biomarker prediction.
[0111] The image discriminator 314 receives a pathological tissue image normalized by the normalization process 310, examines the image including image metadata, and determines the type of the image. The image discriminator 314 can analyze the image data to determine whether the image is a training image, for example, an image from a training dataset. The image discriminator 314 can analyze the image data to determine the type of labeling on the image, for example, whether the image has tile-level labeling, slide-level labeling, or no labeling. The image discriminator 314 can analyze the image data to determine the slide staining used to generate digital images, H&E, IHC, etc.
[0112] In response to examining this image data, the image discriminator 314 determines which images are provided to the slide-level label pipeline 313 for supply to the single-scale classifier module 307 of the deep learning framework, and which images are provided to the tile-level label pipeline 315 for supply to the multi-scale classifier 306 of the deep learning framework.
[0113] In the illustrated example, the pipeline 315 that has tile-level labeling for the images includes a tissue detection process and an image tiling process. These processes may be performed on all received image data, only on the training image data, only on the image data received for analysis, or on some combination thereof. In some examples, for instance, the tissue detection process can be excluded to facilitate deep learning training, image analysis, and / or biomarker prediction. In fact, any of the processes of the controller 302 may be performed in a dedicated biomarker prediction system or distributed for performance by an externally connected system. For example, the pathology imaging system may be configured to perform a normalization process before sending the image data to the biomarker prediction system. In some examples, the biomarker prediction system can communicate an executable normalization software package to a connected external system, and the external system configures those systems to perform normalization or other preprocessing.
[0114] In an example where the image discriminator 314 sends unlabeled images to the pipeline 315, the pipeline 315 includes a multi-instance learning (MIL) controller further described herein, and the MIL controller is configured to convert all or some of these pathology images into tile-labeled images. The MIL controller can be configured to perform the processes as described in FIGS. 18-26 herein.
[0115] To facilitate tissue detection by the trained tissue classifier, the tissue detection process of pipeline 315 can perform an initial tissue identification to find and segment the tissue region of interest for biomarker analysis. Such tissue identification of the tissue region of interest may include, for example, identifying tissue boundaries and segmenting the image into tissue and non-tissue regions, and as a result, the metadata identifying the tissue region is saved together with the image data to facilitate processing and prevent attempts at biomarker analysis in non-tissue regions or regions not corresponding to the tissue under examination.
[0116] To facilitate deep learning classification in various multi-scale configurations, the deep learning framework multi-scale classifier module 306 is configured to classify tissues using tiling analysis. For example, in pipeline 315, the tissue detection process sends a pathological tissue image (e.g., image data enhanced with tissue detection metadata) to an image tiling process, which selects and applies a tiling mask to the received image to break the image into small sub-images for analysis by framework module 306. Pipeline 315 stores multiple different tiling masks and can select a tiling mask. In some examples, the image tiling process selects one or more tiling masks optimized for different biomarkers. That is, in some examples, image tiling is specific to a biomarker. This allows the use of tiles of various pixel sizes and various pixel shapes that are specially selected, for example, to improve the accuracy associated with a particular biomarker and / or to reduce processing time. For example, the tile size optimized for identifying the presence of TILs in an image may be different from the tile size optimized for identifying PD-L1 or another biomarker. Thus, in some examples, preprocessor controller 302 is configured to perform image processing and tiling specific to a type of biomarker, and after the system 300 analyzes the image data of that biomarker, controller 302 can reprocess the original image data for analysis for the next biomarker until all biomarkers have been examined.
[0117] Generally speaking, the efficiency of the operation of deep learning framework module 306 can be increased by selecting the tiling mask applied by the image tiling process of pipeline 315. The tiling mask can be selected based on the size of the received image data, based on the configuration of deep learning framework 306, based on the configuration of framework module 304, or based on some combination thereof.
[0118] The tiling mask may have different sizes of tiling blocks. Some tiling masks have uniform (i.e., each of the same size) tiling blocks. Some tiling masks have tiling blocks of different sizes. The tiling mask applied by the image tiling process can be selected, for example, based on the number of classification layers within the deep learning framework 306. In some examples, the tiling mask can be selected based on the processor configuration of the biomarker prediction system, for example, when multiple parallel processors are available, or when a graphical processing unit or tensor processing unit is used.
[0119] In the illustrated example, the deep learning multi-scale classifier module 304 is configured to perform cell segmentation via the cell segmentation model 316. Here, the cell segmentation can be a pixel-level process of the pathological tissue image from the normalization process 310. In other examples, this pixel-level process may be performed on the image tiles received from the pipeline 315. In some examples, since some of the biomarkers identified herein are determined from cell-level analysis as opposed to tissue-level analysis, the cell segmentation process of the framework 304 provides biomarker classification. These include, for example, ring-shaped nuclei, large nuclei, and high NC ratios. Module 304 can be configured using a CNN configuration, in particular, using an FCN configuration for implementing each separate segmentation.
[0120] The deep learning framework multi-scale classifier module 306 includes a tissue segmentation model 318, a tissue classification model 320, and a biomarker classification model 320. Similar to module 304, module 306 can be configured using a CNN configuration, in particular, using an FCN configuration for implementing each separate segmentation.
[0121] In one example, the cell segmentation model 316 of module 304 is developed by modifying a UNet classifier that forms a three-class segmentation model by replacing the loss function with a cross-entropy function, a focal loss function, or a mean squared error function, and can be configured as a three-class semantic segmentation FCN model. The three-class nature of the FCN model means that the cell segmentation model 316 can be configured as a first pixel-level FCN model that identifies each pixel of the image data and assigns it to a cell sub-unit class, namely (i) inside the cell, (ii) at the cell boundary, or (iii) outside the cell. This is provided as an example. The segmentation size of the module model 316 can be determined based on the type of cell to be segmented. For example, for both TIL biomarkers, the model 316 can be configured to perform lymphocyte identification and segmentation using a three-class FCN model. For example, the cell segmentation model 316 can be configured to classify the pixels in the image as corresponding to (i) the inside, (ii) the boundary, or (iii) the outside of the lymphocyte cell. The cell segmentation model 316 can be configured to identify and segment any number of cells, examples of which include tumor positive, tumor negative, lymphocyte positive, lymphocyte negative, immune cells including lymphocytes, cytotoxic T cells, B cells, NK cells, macrophages, and the like.
[0122] In some examples, module 304 receives the tiled sub-images from pipeline 315, and the cell segmentation model 316 determines a list of the positions of all lymphocytes, and those positions are compared with the list of all cells of the other three class models determined from model 316 to eliminate misdetected lymphocytes that are not cells. Next, the system 300 obtains a new list of the positions of the lymphocytes confirmed from this module 304 and compares it with the list of the tissue segmenter module 318 of the tissue, for example, the positions of the tumor and non-tumor tissues determined from the tissue classification model 320, to determine whether the lymphocytes are in the tumor or non-tumor region.
[0123] Using the three-class model makes it easier to count individual cells and enables more accurate classification, especially when two or more cells overlap with each other. Tumor-infiltrating lymphocytes overlap with tumor cells. In the conventional two-class cell contour model that labels only whether the outer edge of a cell is included in a pixel, each cluster of two or more overlapping cells is counted as one cell.
[0124] In addition to using the three-class model, the cell segmentation model 316 can be configured to avoid the possibility that cells spanning two tiles are counted twice by adding a buffer slightly wider than an average cell around all four sides of each tile. The intention is to count only the cells displayed in the unbuffered area in the center of each tile. In this case, the tiles are arranged such that the unbuffered areas in the centers of adjacent tiles are adjacent and do not overlap. Adjacent tiles overlap in their respective buffer areas.
[0125] In one example, the cell segmentation algorithm of model 316 can be formed from two UNet models. One UNet model can be trained using an image with a mixture of tissue classes in which a human analyst has highlighted the outer edges of each cell and classified each cell according to the tissue class. In one example, the training data includes digital slide images in which all pixels are labeled as either inside a cell, at the outer edge of a cell, or in the background which is outside all cells. In another example, the training data includes digital slide images in which all pixels are labeled with a "yes" or "no" label to indicate whether they show the outer edge of a cell. This UNet model can recognize the outer edges of many types of cells and classify each cell according to the cell shape or its position within the tissue class region assigned by the tissue classification module 320.
[0126] Another UNet model can be trained on images of many cells of a single tissue class, or images of a diverse set of cells where only the cells of one tissue class are outlined in a binary mask. In one example, a training set is labeled by associating a first value with all pixels representing the target cell type and a second value with all other pixels. Visually, an image labeled in this way is displayed as a black and white image, where all pixels representing the target tissue class are white and all other pixels are black, and vice versa. For example, the image can have only labeled lymphocytes. This UNet model can recognize the outer edges of that particular cell type and assign labels to cells of that type within the digital image of the slide.
[0127] The cell segmentation model 316 is a trained cell segmentation model that can be used for cell detection. In some examples, however, the model 316 is configured as a biomarker detection model that is a pixel-level classifier that classifies pixels as corresponding to a biomarker.
[0128] Turning to the multi-scale classifier module 306 of the deep learning framework, the tissue segmentation model 318 can be configured in a similar manner to the segmentation model 316, i.e., by modifying the UNet classifier that forms a three-class segmentation model by replacing the loss function with a cross-entropy function, a focal loss function, or a mean squared error function, and can be configured as a three-class semantic segmentation FCN model. The model 318 can identify the interior, exterior, and boundaries of various tissue types within the tile.
[0129] The tissue classification model 320 is a tile-based classifier configured to classify tiles as corresponding to one of a plurality of different tissue classifications. Examples of tissue classes include, but are not limited to, tumor, stroma, normal, lymphocyte, fat, muscle, blood vessel, immune cluster, necrosis, hyperplasia / dysplasia, erythrocyte, and tissue classes or cell types that are positive (specifically, containing an amount of the IHC staining target molecule greater than a particular threshold) or negative (not containing the molecule or containing an amount of the molecule less than a particular threshold) for an IHC staining target molecule. Examples also include tumor positive, tumor negative, lymphocyte positive, and lymphocyte negative.
[0130] Based on cell segmentation in the pathological tissue image generated by the cell segmentation model 316 and tissue classification from the tissue classification model 302, the biomarker classification model 322 receives data from both and determines the presence of a predicted biomarker in the pathological tissue image, and in particular, in a multi-scale configuration, determines the presence of a predicted biomarker in each tile image of the pathological tissue image. The biomarker classification model 322 can be a trained classifier implemented in the deep learning framework model 306 as shown, or implemented separately from the model 306 such as the deep learning post-processing controller 308.
[0131] In some examples of the biomarker classification model that detects TIL biomarkers, the tissue classification model 320 is trained to identify the percentage of TILs within a tile image, the cell segmenter 316 determines cell boundaries, and the biomarker classification model 322 classifies the tile image based on the percentage of TILs within the cell interior, resulting in a classification of (i) tumor-IHC / lymphocyte positive or (ii) non-tumor-IHC / lymphocyte positive.
[0132] In some examples of the biomarker classification model for detecting ploidy, the biomarker classification model 322 can be trained using, for example, the technique provided in "Classification and mutation prediction from non-small cell lung cancer histopathology images using deep learning" by Coudray N, Ocampo PS, Sakellaropoulos T, Narula N, Snuderl M, Fenyo D, et al., Nat Med. 2018;24:1559-67, based on histopathological images and related ploidy scores.
[0133] In one example, the training data can be data such as the actual karyotype determined by a cytogeneticist. In some examples, however, the biomarker classification model 322 can be configured to infer such data. Ploidy data can be formatted in columns of chromosome number, start position, stop position, and length of the region. The ploidy score can be determined from DNA sequencing data and can be specific to a gene, chromosome, or chromosomal arm. The ploidy score can be global and may represent the entire genome of the sample (global CNV / copy number variation can cause changes in hematoxylin staining of tumor nuclei), and can be a score calculated by averaging the ploidy scores of each region within the genome. The local regional ploidy score can be weighted according to the length of each region associated with that score. The trained ploidy model of the biomarker classification model 322 can be specific to a gene, chromosomal arm, or entire chromosome. This is because each section may have a different effect on the cell morphology seen in the histopathological image. Even if the tumor purity or cell count on the slide is low, if the ploidy is higher than normal, there may be sufficient material left for genetic testing, so the predicted ploidy biomarker data can affect the acceptance / rejection analysis. The biomarker metric processor 326 can be configured to make such a determination prior to report generation.
[0134] In some examples of the biomarker classification model for detecting the signet ring morphology, the biomarker classification model 322 can be trained with a signet ring morphology model based on classification techniques such as low cohesion (PC), signet ring cells (SRC), and Lauren's sub-classification, as well as Mariette, C., Carneiro, F., Grabsch, H. I. et al.'s "Consensus on the pathological definition and classification of poorly cohesive gastric carcinoma" Gastric Cancer 22, 1-9 (2019) and other signet ring morphology classifications.
[0135] In some examples of the biomarker classification model for detecting the NC ratio, the cell segmentation model 316 can be composed of the 3-class UNet described herein, but this model is trained to distinguish three classes, namely the nucleus, cytoplasm, and cell boundary / non-cell background. In one example, the training data may be an image in which each pixel is manually annotated with one of these three classes, and / or an image thus annotated by the trained model as illustrated in the example of the updated training image of FIG. 4.
[0136] Thus, the cell segmentation model 316 can be trained to analyze the input image, assign one of the three classes to each pixel, and define the cell as a group of all cytoplasmic pixels between the adjacent nuclear pixels and the next closest boundary pixels. Next, the biomarker classification model 322 can be configured to calculate the nucleus:cytoplasm ratio for each cell as the area (number of pixels) of the cell's nucleus divided by the area (number of pixels) of the entire cell (nucleus and cytoplasm).
[0137] To identify tumor tissue and the status of tumors in tissue, in one example, the deep learning framework 306 can be configured using a FCN classifier. In one example, the deep learning framework 304 can be configured as a pixel-resolution FCN classifier, while the deep learning framework 306 can be configured as a tile-resolution FCN classification model, or a tile-resolution CNN model, i.e., a model that performs classification on the entire received tile of image data.
[0138] The classification model 320 of module 306 can be configured to classify the tissue within a tile as corresponding to one of a plurality of tissue classes, such as, for example, the status of a biomarker, the status of a tumor, the tissue type, and / or the state / condition of a tumor, or other information. In the illustrated example, module 306 is configured to have a tissue classification 320 and a tissue segmentation model 322. In an exemplary implementation of the TIL biomarker, the tissue classification model 320 can classify the tissue using tissue classifications such as tumor-IHC positive, tumor-IHC negative, necrosis, stroma, epithelium, or blood. The tissue segmentation model 328 identifies the boundaries of different tissue types identified by the tissue classification model 320 and generates metadata for use by the post-processing controller 308 to visually display the boundaries and color-coding of different tissue types in the overlay mapping report generator.
[0139] In an exemplary implementation, the deep learning framework 300 performs tile-based classification by receiving tiles (i.e., sub-images) from the processes of the classification models 320 and 322. In some examples, tiling can be performed by the framework 306 using a tiling mask, and the module 306 itself can send the sub-images generated for pixel-level segmentation to the module 304 in addition to performing tissue classification. The module 306 can inspect each tile sequentially one after another, or the module 306 can inspect each tile in parallel depending on the nature of the matrix generated by the FCN model for faster processing of the image.
[0140] In some examples, the tissue segmentation model 318 receives pixel-resolution cell segmentation data and / or pixel-resolution biomarker segmentation data from the module 304 and performs statistical analysis on a tile-by-tile and basis. In some examples, the statistical analysis determines (i) the area of the image data covered by the tissue, e.g., the area of a stained pathological tissue slide covered by the tissue, and (ii) the number of cells in the image data, e.g., the number of cells in a stained pathological tissue slide. For example, the tissue segmentation model 318 can accumulate the classification of cells and tissues for each tile of the image until all the tiles forming the image are classified.
[0141] When the deep learning framework is a multi-scale classifier module for classifying biomarkers in a tile-based manner, the deep learning framework 300 is further configured to classify biomarkers using the classifications trained from the slide-level training images without requiring tile-level labeling. For example, as further discussed below, the slide-level training images received by the image discriminator can be provided to a slide-level label pipeliner 313 having a MIL controller, and the MIL controller executes a process as described in FIGS. 18-26 of this specification to generate a plurality of tile images having the inferred classifications, and optionally performs tile selection on those tiles, and is configured to train an organization classification model 317 and a biomarker classification model 319. Examples of single-scale classifiers include a CMS biomarker classification model whose output is the CMS class, and an HRD biomarker classification model whose output is HRD+ or HRD-. These classifications may be performed on the entire pathological tissue image to determine biomarker predictions, may be performed on each tile image of the digital image, and may be analyzed by a biomarker metric processor 326 to determine biomarker predictions from the tile images.
[0142] In some examples of a biomarker classification model for detecting HRD, the biomarker classification model 319 can be configured to predict HRD. Training of the HRD model within the classifier 318 can be based on a pathological tissue image and a matched HRD score. For example, the training data can be generated by H&E images and RNA sequence data, which in some examples includes RNA expression profile data, and the RNA expression profile data is supplied to the HRD model and fed back for further training, similar to the updated training data 403 in FIG. 4. The training data can be derived from tumor organoids, and the H&E image of the organoid is paired with a measurement of the sensitivity of the organoid to a PARP inhibitor indicating HRD, or the result of an HRD model run by the RNA team based on the RNA expression profile of the organoid. Exemplary HRD prediction models of the biomarker classification model 319 are described in Peng, Guang et al., "Genome-wide transcriptome profiling of homologous recombination DNA repair", Nature communications vol. 5 (2014): 3361 and van Laar, R.K., Ma, X.-J., de Jong, D., Wehkamp, D., Floore, A.N., Warmoes, M.O., Simon, I., Wang, W., Erlander, M., van’t Veer, L.J. and Glas, A.M. (2009), "Implementation of a novel microarray-based diagnostic test for cancer of unknown primary", Int. J. Cancer, 125: 1390-1397.
[0143] In one example, a deep learning framework can use RNA expression to identify HRD from H&E slides and identify slide-level labels indicating the percentage of slides containing biomarker-expressing cells. In one example, an activation map approach for RNA labels can be applied across the whole slide as a binary label (i.e., positive or negative HRD expression somewhere in the tissue) or as a continuous percentage (i.e., if it is found that 62% of the cells in the image express HRD). Binary RNA labels can be generated by next-generation sequencing of the specimen, and cell percentage labels can be generated by applying single-cell RNA sequencing. In one example, single-cell sequencing can identify the cell types and amounts present in the RNA expression from NGS.
[0144] Training of a tile-based deep learning network to predict biomarker classification labels for each tile of a whole-slide image can be performed using any of the methods described herein. Once training is complete, the model can be applied to the method of activation mapping to each tile. Activation mapping can be performed using Grad-CAM (Gradient Class Activation Mapping) or guided backpropagation. Both can identify which regions of the tile contribute most to the classification. In one example, the portion of the tile that contributes most to the HRD positive class can be cells clustered in the upper right corner of the tile. Next, the cells within the identified active region can be labeled as HRD positive cells.
[0145] To prove that the model clinically functions reliably, comparing the results of the model to a ground truth source may be involved. One possible method of generating the ground truth may involve separating small regions of tissue, each containing fewer than 100 cells, by segmenting via tissue microarray and sequencing each region individually to obtain RNA labels for each region. The method of generating the ground truth may further involve using a biomarker classification model to classify these regions and determining the accuracy with which the activation map highlights cells in regions of high HRD expression and ignores most cells in regions of low HRD expression.
[0146] When training a tile-based deep learning network to predict biomarker classification labels for each tile, a strong teacher approach is utilized to generate the biomarker labels and identify the HRD status (positive or negative) of individual cells. Single-cell RNA sequencing can be used alone or in combination with laser-induced microdissection to extract one cell at a time and create a label for each cell. In one example, a cell segmentation model can be incorporated to first obtain the cell outline, and then an artificial intelligence engine can be incorporated to classify the pixel values within the outline of each cell according to the biomarker status. In another example, a mask of the image can be generated, with a first value assigned to HRD-positive cells and a second value assigned to HRD-negative cells. Next, a single-scale deep learning framework can be trained using the masked slide to identify cells that express HRD.
[0147] In some examples of a biomarker classification model for detecting CMS, the biomarker classification model 319 can be configured to predict CMS. Such biomarker classification can be configured to classify segmented cells as corresponding to cancer-specific classifications. For example, the four trained CMS classifications in primary colorectal cancer include 1-immune infiltration (often BRAFmut, MSI-High, TMB-High), 2-normal (often ERBB / MYC / WNT-driven), 3-metabolic (often KRASmut), and 4-mesenchymal (often TGF-B-driven). In other examples, more trained CMS classifications can be used, but generally, two or more CMS subtypes are classified in the examples herein. Further, other cancer types may have their own trained CMS categories, and the classifier 318 can be configured to have a model for subtyping each cancer type. Technical examples for developing CMS classifications of 4, 5, 6, 7 or more numbers of CMS classifications are described in Eide, P.W., Bruun, J., Lothe, R.A. et al., "CMScaller: an R package for consensus molecular subtyping of colorectal cancer pre-clinical models", Sci Rep 7, 16618 (2017) and https: / / github.com / peterawe / CMScaller.
[0148] Training of the CMS model within classifier 318 can be obtained based on pathological tissue images that match the CMS category assignment. The assignment of the CMS category can be based on the RNA expression profile. In one example, it is generated by an R program that uses the nearest template prediction, referred to as the CMS Caller (see "CMScaller: an R package for consensus molecular subtyping of colorectal cancer pre - clinical models" by Eide, P.W., Bruun, J., Lothe, R.A. et al., Sci Rep 7, 16618 (2017) and https: / / github.com / peterawe / CMScaller). An alternative classification using a random forest model is described in "The consensus molecular subtypes of colorectal cancer" by Guinney, J., Dienstmann, R., Wang, X. et al., Nat Med 21, 1350 - 1356 (2015). For example, the CMS Caller examines each RNA sequence data sample to determine whether each gene is above or below the mean, and a binary classification of each gene is performed. This avoids, for example, batch effects between different RNA sequence datasets. The training data can also include DNA data, IHC data, mucin markers, treatment response / survival data from clinical reports. These may or may not be associated with the assignment of the CMS category. For example, CMS4 IHC stains positive for TGFbeta, CMS1 IHC may be positive for CD3 / CD8, CMS2 and 3 have changes in mucin genes, CMS2 responds to cetuximab, and CMS1 responds well to bevacizumab. The survival prognosis of CMS1 is the best, and that of CMS4 is the worst. (See slide 12 of the CMS slides). CMS categories 1 and 4 can be detected from H&E. By performing training, for example, the model can be trained using the architecture of FIG. 4 to identify and classify the differences between CMS2 and 3.
[0149] In one example, the biomarker classification model 319 can be configured to identify five CRC-specific subtypes (CRIS) with distinct molecular, functional, and phenotypic specificities, namely (i) CRIS-A: mucinous, glycolytic, microsatellite instability-high or KRAS mutation, (ii) CRIS-B: TGF-β pathway activity, epithelial-mesenchymal transition, poor prognosis, (iii) CRIS-C: elevated EGFR signaling, sensitivity to EGFR inhibitors, (iv) CRIS-D: WNT activation, IGF2 gene overexpression and amplification, and (v) CRIS-E: phenotype like Paneth cells, TP53 mutation. The CRIS subtypes can correctly classify independent sets of primary and metastatic CRCs, but the overlap with existing transcriptional classes is limited, and the prediction and prognostic performance was unprecedented. See, for example, Isella, C., Brundu, F., Bellomo, S. et al., “Selective analysis of cancer-cell intrinsic transcriptional traits defines novel clinically relevant subtypes of colorectal cancer” Nat Commun 8, 15107 (2017).
[0150] In the case of biomarker detection, instead of attempting a mere average CMS classification across all tiles, the biomarker classification model 319 can be trained with a CMS model that predicts for each tile and a CMS classification that identifies different tissue types (e.g., stroma) for that classification. In one example, each tile is processed and the CMS model uses the pixel data associated with each tile to generate a compressed representation, and each tile is assigned to a class (cluster 1, cluster 2, etc.) based on the pattern of pixel data for each tile and the similarity between tiles. A list of the percentage of tiles within an image belonging to each cluster is the cluster profile of the image and can be provided by a report generator. In one example, each profile is fed into the model along with the corresponding CMS designation or RNA expression profile (which is the original method used to define the CMS categories) for training. In another example, each tile of all training slide images is annotated according to the overall CMS category assigned to the entire slide from which that tile originated, and the tiles are clustered and analyzed to determine the cluster most closely associated with the CMS category.
[0151] In some examples, the biomarker classification model 319 (and biomarker classification model 322) can cluster each tile into a discrete number of clusters instead of performing the same classification on each tile and equally weighting that tile classification. One way to achieve this is to include an attention layer in the biomarker classification model. In one example, all tiles of all training slides can be classified into clusters, and then if the number of tiles within a cluster is not statistically related to the biomarker, that cluster is not weighted as highly as clusters associated with the biomarker. In other examples, majority voting techniques can be used to train the biomarker classification model 319 (or model 322).
[0152] Although shown as separate models, biomarker classification models 322 and 319 can each be configured to include all or part of the corresponding tissue classification model, cell segmentation model, and tissue segmentation model, as in the case of classifying various biomarkers herein. Further, biomarker classification model 322 is shown to be included within multi-scale classifier module 306, and biomarker classification model 319 is shown to be included within single-scale classifier module 307, but in some examples, all or part of these biomarker classification models can be implemented in post-processing controller 308, as in the case of classifying various biomarkers herein. Further, although described as a tile-level or slide-level classification model, in some examples, biomarker classification models 322 and 319 can be configured as pixel-level classifiers in some examples.
[0153] From the decisions made by modules 304, 306, and 307, post-processing controller 308 can determine whether the image data exceeds a threshold and / or includes an amount of tissue that meets a criterion, e.g., sufficient tissue for genetic analysis, sufficient tissue to use the image data as training images in the training phase of a deep learning framework, or sufficient tissue to combine the image data with an existing trained classifier model.
[0154] Accordingly, patient reports can be generated in various examples herein, including those described with reference to FIG. 3 and elsewhere herein. This report can be presented to the patient, physician, healthcare provider, or researcher in a digital copy (e.g., JSON object, PDF file, or image on a website or portal), hard copy (e.g., printout on paper or another tangible medium), audio (e.g., recording or streaming), or another format.
[0155] The report may include information related to gene expression calls (e.g., overexpression or underexpression of a particular gene), detected genetic mutations, other characteristics of the patient's sample, and / or clinical records. The report may further include clinical trials for which the patient is eligible, treatment methods that may be applicable to the patient, and / or side effects predicted if the patient undergoes a particular treatment, based on the detected genetic mutations, other characteristics of the sample, and / or clinical records.
[0156] Using the results included in the report and / or additional results (e.g., from a bioinformatics pipeline), a database of clinical data can be analyzed to determine, in particular, whether the treatment delays the progression of cancer in other patients with the same or similar results as the specimen. The results can also be used in the design of tumor organoid experiments. For example, the organoids can be genetically engineered to have the same characteristics as the specimen and, after being subjected to treatment, observed to determine whether the treatment can reduce the growth rate of the organoids and thus potentially reduce the growth rate of the patient associated with the specimen.
[0157] In one example, the post - processor 308 is further configured to determine a plurality of different biomarker prediction metrics and a plurality of tumor prediction metrics, for example, using the biomarker metric processing module 326. Examples of prediction metrics include tumor purity, the number of tiles classified as a particular tissue class, the number of cells, the number of tumor - infiltrating lymphocytes, clustering of cell types or tissue classes, density of cell types or tissue classes, characteristics of tumor cells (roundness, length, nuclear density), thickness of the stroma around the tumor tissue, image pixel data statistics, predicted patient survival, PD - L1 status, MSI, TMB, tumor origin, and immunotherapy / treatment response.
[0158] For example, the biomarker metric processing module 326 can determine, for each tissue class, the number of tiles classified into one or more single tissue classes, the percentage of tiles classified into each tissue class, for any two classes, the ratio of the number of tiles classified into the first tissue class to the number of tiles classified into the second tissue class, and / or the total area of tiles classified into a single tissue class. The module 326 can determine tumor purity based on the number of tiles classified as tumor relative to other tissue classes, or based on the number of cells located in tumor tiles relative to the number of cells located in other tissue class tiles. The module 326 can determine the number of cells for the entire pathological tissue image, within a region pre-defined by the user, within a tile classified as any of the tissue classes, within a single grid tile, or over the entire region or area of interest, regardless of whether it is pre-determined, selected by the user during operation of the system 300, or automatically selected by the system 300, for example, by determining the most likely region of interest based on image analysis. The module 326 can determine clustering of cell types of the tissue class based on the spacing and density of the classified cells, the spacing and distance of the tiles classified into the tissue class, or any visually detectable feature. In some examples, the module 326 determines the probability that two adjacent cells are, for example, either two immune cells, two tumor cells, or one of each. The module 326 determines the characteristics of the tumor cells by determining the average circularity, perimeter length, and / or nuclear density of the identified tumor cells. The identified stromal thickness can be used as a predictor of the patient's response to treatment. The image pixel data statistics determined by the module 326 can include the average, standard deviation, and sum of each tile of any single image or collection of images of pixel data, including red, green, blue (RGB) values, optical density, hue, saturation, grayscale, and deconvolution of staining.Furthermore, module 326 can calculate the position of lines, the pattern of alternating brightness, the contour of the shape, the segmented tissue classes and / or the staining pattern of the segmented cells in the image. In any of these examples, module 326 can be configured to create a determined / predicted status, and then the overlay display generation module 324 generates a report for displaying the determined information. For example, the overlay map generation module 324 can generate a network-accessible user interface, whereby the user can select different types of data to be displayed. Module 324 generates an overlay map showing the selected different types of data overlaid on the rendition of the original stained image data.
[0159] FIG. 4 shows a machine learning data input / flow schematic 400 that can be implemented in the system 300 of FIG. 3 or, more generally, in any of the systems and processes described herein.
[0160] In the training mode in which the deep learning framework of system 300 is trained, various training data can be obtained. In the illustrated example, training image data 401 in the form of high-resolution and low-resolution pathological tissue images is provided to preprocessing controller 302. As shown, the training images can include annotated tissue image data from various tissue types, such as tumors, stroma, normal, immune clusters, necrosis, hyperplasia / dysplasia, and erythrocytes. As shown, the training images can include computer-generated synthetic image data, as well as image data of segmented cells (cell image data) and slide-level labels or tile-level labels (collectively, image data labeled with biomarkers), such as the biomarkers discussed herein. These training images can be digitally annotated, but in some examples, the annotation of the tissue is done manually. In some images, the training image data includes molecular data and / or demographic data, for example, as metadata within the image data. In the illustrated example, such data is separately supplied to deep learning framework 402 (consisting of an exemplary implementation of multi-scale deep learning framework 306’ and single-scale deep learning framework 307’). For additional training of the deep learning framework, other training data, such as pathway activation scores, can also be provided to controller 302.
[0161] In some examples, deep learning framework 402 generates updated training images 403, and the updated training images 403 are annotated and segmented by deep learning framework 402 and fed back to framework 402 (or preprocessing controller 302) for use in updated training of the framework.
[0162] In the diagnostic mode, patient image data 405 is provided to controller 302 and used according to the examples herein.
[0163] Any of the image data in this specification, including patient portrait data and training images, can be pathological tissue image data such as H&E slide images and / or IHC slide images. For example, in the case of IHC training images, the images may be segmented images that distinguish cytotoxic T cells from regulatory T cells, or other cell types.
[0164] In some examples, the controller 302 generates image tiles 407 and accesses one or more tiling masks 409 and tile metadata 411. These are supplied as inputs to the deep learning framework 402 for the controller 302 to determine the predicted biomarkers and / or the status and metrics of the tumor, and then provided to the overlay report generator 404 for generating the biomarker and tumor report 406. Optionally, the report 406 may include an overlay of the pathological tissue image and, in one example, may further include biomarker scoring data such as percentage TIL.
[0165] In some examples, the clinical data 413 is provided to the deep learning framework 402 for use in the analysis of the image data. The clinical data 413 may include health records, biopsy tissue type, and anatomical location of the biopsy. In some examples, the tumor response data 415 collected from the patient after treatment is additionally provided to the deep learning framework 402 to determine changes in the status of the biomarkers, the status of the tumor, and / or their metrics.
[0166] Figure 5 shows an exemplary deep learning framework 500 formed from a plurality of different biomarker classification models. The elements of Figure 5 are provided as follows. "Cell" refers to a cell segmentation model, e.g., a trained pixel-level segmentation model, according to an example of this specification. "Multi" refers to a multi-scale (tile-based) tissue classification model according to an example of this specification. "Post" refers to arithmetic calculations that can be performed at the final stage of a biomarker classification model configured to predict the biomarker status of an image or tile image according to an example of this specification in response to one or more data from the "Cell" or "Multi" stage. In one example, "Post" may include a majority vote, e.g., identifying to sum the number of tiles in an image associated with each biomarker label and assign the biomarker status in the image to the biomarker label with the largest sum. The two-layer "Post" configuration refers to a two-stage post-processing configuration where arithmetic calculations can be stacked. In one example, the first layer of Post may include summing the tissue and the cells in the labeled tiles and summing the lymphocyte cells in the same tiles. The second layer can generate a ratio by dividing the number of lymphocyte cells by the number of cells and use this to assign the status of the biomarker in the image based on whether the ratio exceeds a threshold when compared to the threshold. The final "Post" configuration may include other post-processing functions such as the report generation process described in this specification. "Single" refers to a single-scale classification model according to an example of this specification. "MIL" refers to an MIL controller according to an example of this specification. In the illustrated example, the deep learning framework 500 includes a TIL biomarker classification model 502, a PD-L1 classification model 504, a first CMS classification model 506 based on a "Single" classification architecture, and a second CMS classification model 508 based on a "Multi" classification architecture, as well as an HRD classification model 510. Patient data 512 such as molecular data, demographic data, tumor response data, and patient images 514 are stored in a dataset accessible by the deep learning framework 500.
[0167] The training data is also shown in the form of cell segmentation training data 516, single-scale classification biomarker training data 518, multi-scale classification biomarker training data 520, MIL training data 522, and post-processing training data 524.
[0168] FIG. 6 shows a process 600 that can be executed by an imaging-based biomarker prediction system 102, a deep learning framework 300, or a deep learning framework 402, particularly in a deep learning framework having a multi-scale configuration.
[0169] As part of the training process, at block 602, a pathological tissue image with tile labels is received by the deep learning framework 300. Here, the pathological tissue image can be of any type, but in this example, it is shown as a digital H&E slide image. These images can be (for example, in the case of a supervised learning configuration) training images of cancer types that have been determined and labeled (and thus known) in the past. In some examples, the images can be training images of multiple different cancer types. In some examples, the images can be (for example, in the case of an unsupervised learning configuration) training images that are unknown or include some or all of the images of cancer types without labels. In some examples, the training images include digital H&E slide images annotated with tissue classes (for training a tile-resolution FCN tissue classifier) and other digital H&E slide images annotated with each cell. In an example of training for TIL biomarker classification, each lymphocyte can be annotated with an H&E slide image (for example, to train an annotated image of a pixel-resolution FCN segmentation classifier to train a UNet model classifier). In some examples, the training images can be digital IHC staining images, particularly images where the IHC staining targets lymphocyte markers, for training a pixel-resolution FCN segmentation classifier. In some examples, the training images include images combined with molecular data, clinical data, or other annotations (such as pathway activation scores).
[0170] In the illustrated example, at block 604, preprocessing is performed on the training images, such as the normalization process described herein. Other preprocessing processes described herein may also be performed at block 604.
[0171] At block 606, the tile-labeled H&E slide training images are provided to a deep learning framework and analyzed within a machine learning configuration, such as a CNN, and more specifically, in some examples implemented as an FCN model, a tile image of the training image for tissue classification training, a pixel of the training image for cell segmentation training, and in some examples a tile image for biomarker classification training. As a result, at block 608, a trained deep learning framework multi-scale biomarker classification model is generated, which may include a cell segmentation model and a tissue classification model. When training multiple biomarker classification models, at block 608, separate models can be generated for each of the biomarkers TIL, PD-L1, ploidy, NC ratio, and signet ring morphology.
[0172] As a prediction process, at block 610, a new unlabeled pathological tissue image, such as an H&E slide image, is received and provided to the multi-scale biomarker classification model, and at block 612, the biomarker status of the received pathological tissue image is predicted as determined by one or more biomarker classification models.
[0173] For example, at block 610, a new (unlabeled or labeled) pathological tissue image can be received from a physical clinical record system or a primary care system and applied to a trained deep learning framework that applies its trained cell segmentation, tissue classification model, and biomarker classification model. At block 612, a biomarker prediction score is determined. The prediction score can be determined for the entire pathological tissue image or various regions of the entire image. For example, for each image, at block 612, the absolute number of biomarkers on the image, the percentage of the number of cells within the tumor region associated with each biomarker, and / or the classification of the biomarker or specification of other information can be generated. In some examples, the deep learning framework may identify the predicted biomarkers for all tissue classes identified within the image. As such, the biomarker prediction probability scores may vary across the entire image. For example, when predicting the presence of TILs, process 612 may predict the presence of TILs at different locations within the pathological tissue image. As a result, the prediction of TILs varies across the entire image. This is provided as an example, and at block 612, any number of metrics discussed herein can be determined.
[0174] As shown in process 900 of FIG. 9, at block 902, after the prediction is made, the predicted biomarker classification can be received by a report generator. At block 904, a pathological tissue image, and thus a clinical report for the patient, can be generated, which includes the status of the predicted biomarkers. At block 906, an overlay map indicating the predicted biomarker status can be generated, which is provided to a clinician for display or to a pathologist to determine a preferred immunotherapy corresponding to the predicted biomarker.
[0175] In FIG. 7, an exemplary process 700 is provided for determining the status of a predicted biomarker, particularly for predicting the status of TIL. The process 700 can still be used to predict the status of any number of biomarkers and other metrics according to the examples described herein.
[0176] The preprocessing controller receives a pathological tissue image and performs initial image processing (702), as described herein. In one example, the deep learning preprocessing controller receives the entire image file in any pyramidal TIFF format and identifies the edges and contours of the viable tissue within the image (e.g., performs segmentation). The output of block 702 can be, for example, a binary mask of the input image where each pixel has a value of 0 or 1. Here, 0 indicates the background and 1 indicates the foreground / tissue. The dimensions of the mask can be the dimensions of the input slide when downsampled by a factor of 128. This binary mask can be temporarily buffered and provided to the tiling process 704.
[0177] In process 704, the preprocessing controller applies a tissue masking process using a tiling procedure to divide the image into sub-images (i.e., tiles) that are inspected individually. Since the deep learning framework is configured to execute two different learning models (one for tissue classification and the other for cell / lymphocyte segmentation), different tiling procedures for each model can be executed in procedure 704. Process 704 can generate two outputs, each of which includes, for example, a list of coordinates defined from the upper left corner of the tile. The output list is temporarily buffered and can be passed to the tissue classification and cell segmentation processes.
[0178] In the example of FIG. 7, tissue classification is performed in process 706, which receives a pathological tissue image from process 704 and performs tissue classification on each received tile using a trained tissue classification model. The trained tissue classification model is configured to classify each tile into different tissue classes (e.g., tumor, stroma, normal epithelium, etc.). To reduce computational redundancy, process 706 can use multiple tiling layers. The trained tissue classification model calculates the class probability of each class stored in the model for each tile. Next, in process 706, the most likely class is determined and that class is assigned to the tile. Process 706 can, as a result, output a single list out of a plurality of lists. Each nested internal list functions as a nested classification, which describes a single tile and includes each of the tile's position, the probability that the tile is each class included in the model, and the ID of the most likely class. This information is listed for each tile. A single list out of the plurality of lists can be saved to the output json file of the deep learning framework pipeline.
[0179] In processes 708 and 710, cell segmentation and lymphocyte segmentation are respectively performed. In processes 708 and 710, a pathological tissue image and a tile list are received from processes 704 and 706. In process 708, a trained cell segmentation model is applied. In process 710, a trained lymphocyte segmentation model is applied. That is, in the example of the figure, for each tile in the cell segmentation tile list, two pixel resolution models are executed in parallel. In one example, both models use the UNet architecture but are trained with different training data. The cell segmentation model identifies cells and draws a boundary line around all the cells in the received tile. The lymphocyte segmentation model identifies lymphocytes and draws a boundary line around all the lymphocytes in the tile. Since hematoxylin binds to DNA, performing "cell segmentation" using a digital H&E slide image may also be referred to as nuclear segmentation. That is, in the cell segmentation model process 708, nuclear segmentation is performed on all cells, and in the lymphocyte segmentation model process 710, nuclear segmentation is performed on lymphocytes.
[0180] In this example, since the same UNet architecture is used for both, processes 708 and 710 each generate two mask array outputs of the same format. Each output is a mask array of the same shape and size as the received tile. Each array element is either 0, 1, or 2. Here, 0 indicates a pixel / position predicted as the background (i.e., outside the object), 1 indicates a pixel / position predicted as the boundary of the object, and 2 indicates a pixel / position predicted as inside the object. In the case of the output of the cell segmentation model, the object refers to a cell. In the case of the lymphocyte segmentation model, the object refers to a lymphocyte. These mask array outputs can be temporarily buffered and provided to processes 712 and 714 respectively.
[0181] In processes 712 and 714, the output mask array of the cell segmentation (UNet) model and the output mask array of the lymphocyte segmentation (UNet) model are received, respectively. Processes 712 and 714 are executed for each received tile and are used to represent the information in the mask array at coordinates in the coordinate space of the original full-slide image.
[0182] In one example, in process 712, access is made to the stored image processing library and the library is used to find the outer shape around the cell interior class, i.e., the outer shape corresponding to the positions having a value of 2 in each mask. In this way, in process 712, a cell alignment process can be executed. In the cell boundary class (indicated by the places having a value of 1 in each mask), the separation between adjacent cell interiors is ensured. Thereby, a list of all outer shapes of each mask is generated. Next, by treating each outer shape as a filled polygon, in process 712, the coordinates of the centroid (center of mass) of the outer shape are determined, and from there, a centroid list is generated by process 712. Next, the coordinates of each of the outer shape list and the centroid list are shifted to generate an output in a coordinate space defined by the entire received image rather than in a coordinate space specific to a single tile within the image. Without this shift, each coordinate would be in the coordinate space of the image tile that contains it. The value of each shift is the same as the coordinates of the upper left corner of the parent tile of the received image. In this example, process 714 executes the same process as process 712 for the lymphocyte class.
[0183] In processes 712 and 714, an outer contour list output and a centroid list output corresponding to each UNet segmentation model are generated respectively. The outer contour is a set of coordinates that, when connected, depict the contour of the detected object. Each outer contour can be represented as a text line composed of constituent coordinates printed in order as pairs of numbers. Each outer contour list from processes 712 and 714 can be saved as a text file composed of such many lines. The centroid list is a list of pairs of numerical values. Each of these outputs can be temporarily buffered and provided to process 716.
[0184] In process 716, it receives the tissue classification output (a single list among multiple lists) from process 706, the cell centroid and outer contour lists from process 712, and the lymphocyte centroid and outer contour lists from process 714, and performs cell segmentation integration.
[0185] For example, in process 716, it can integrate the paired outputs of processes 712 and 714 to generate a single concise list containing the most important information regarding cells. In one example, there are two main components in process 716.
[0186] In the first component of process 716, the information found in the cell segmentation model and the lymphocyte segmentation model is combined. Before the information is combined, this exists as a list of cell outlines and a list of lymphocyte outlines, but since the lymphocyte outlines are the output of two independent models (712 and 714), they are not necessarily a subset of the cell outlines. Thus, (1) since lymphocytes are a type of cell, an object cannot biologically become a lymphocyte if it is not a cell, and (2) since it is desirable to report the percentage of a dataset with the same denominator, lymphocytes are preferably made a subset of cells. Thus, the cell segmentation integration process 716 can be performed by comparing the position of each cell with the positions of all lymphocytes (in one example, this can only be done for the objects within a single tile, so the number of comparisons does not become excessive). In one example, a cell is considered a lymphocyte only if it is "sufficiently close" to a lymphocyte. The definition of "sufficiently close" can be established by empirically determining the median radius of the objects detected by the lymphocyte segmentation model across a set of training histopathological images. This updated set of training images (e.g., 403 in FIG. 4) is an annotated image where the set of training histopathological images is generated by the model itself, and as a result, there are orders of magnitude more images, e.g., millions of automatically annotated images, which form a new or updated training set, which is different from the training set of images used to train the model. In fact, the training set of the model may continue to grow with new received medical images that meet the acceptance / rejection criteria. This may be the case for tissue classification models, as well as cell segmentation and lymphocyte segmentation models. By generating a new training set from the model and evaluating subsequent images, the model can (1) use the median of millions of cells, not just those outlined in the training tiles, and (2) compare the actual size of the detected objects, not the size of the human-drawn annotations.Since the nuclei of lymphocytes are usually spherical, in one example, these objects are all modeled as circles (since they are two-dimensional slices of spheres). The radii of these circles were calculated and the median value was used to determine the typical size for lymphocyte detection. As a result, the final cell list is exactly the same as the objects detected by the cell segmentation model, but the purpose of the lymphocyte segmentation model is to provide a boolean true / false label for each cell in that list.
[0187] In the second component of process 716, each cell is binned into one of the tissue classification tiles (from process 706) based on its location. Note that in the example described here, due to the different architecture of the model, the size of the cell segmentation tiles can be different from the tissue classification tiles. Nevertheless, process 716 has the coordinates of each cell centroid, the coordinates of the upper left corner of each tissue classification tile, and the size of each tissue classification tile, and is configured to determine the parent tile of each cell based on the location of the centroid.
[0188] Process 716 generates an output that is a single list out of a plurality of lists. Each nested internal list functions as a nested classification that describes a single cell and includes the coordinates of the cell centroid, the tile number of the parent tile, the tissue class of the parent tile, and whether the cell is classified as a lymphocyte. This information is listed for each cell and the output list is saved to the output json file of the deep learning framework pipeline.
[0189] In process 718, which can be implemented by a post-processing controller such as biomarker metric processing module 326, in particular in this example, one of a plurality of different biomarker metrics of the predicted TIL status and other TIL metrics is determined as described.
[0190] For example, process 718 can be configured to perform tissue area calculation based on the tissue mask used in process 704 to determine the area covered by the tissue. In some examples, since the tissue mask is a boolean array that takes a value of 1 when tissue is present and 0 elsewhere, process 718 can count the number of 1s to give a measurement of the tissue area. This value is the number of square pixels in a 128-fold downsampling. Multiplying this by 16384 (i.e., for a 128*128 tissue mask), the number of square pixels at native resolution (referred to as "x") is obtained. The native resolution of the image indicates the number of pixels per micron, and taking the square of this number gives the number of square pixels per square micron (referred to as "y"). Dividing the number of square pixels at native resolution by this resolution scaling factor (or x / y using the variables defined above), the number of square microns covered by the tissue is obtained, and thus the tissue area is calculated. That is, process 718 can generate a floating-point number in [0,∞] representing the tissue area in square microns. This value can be used in the acceptance / rejection model process described below.
[0191] As an example of other biomarker statistics, process 718 can be further configured to perform total nucleus calculation using the cell segmentation integration output from process 716. For example, the total number of nuclei on the slide is determined as the number of entries in the cell segmentation integration output. Also, in process 718, the calculation of the percentage of tumor nuclei can be performed based on the output from this process 716. The total number of tumor nuclei in the image is the number of entries in the cell segmentation integration output that meet the requirements that (i) the tissue class of the parent tile is a tumor and (ii) the cell is not classified as a lymphocyte.
[0192] In addition to determining biomarker statistics, process 718 can be further configured to perform an acceptance / rejection process based on the tissue area, total nucleus count, and tumor nucleus count sought. In one example, process 718 can be configured with a logistic regression model, and these three variables are used as inputs, and the output of the model is a binary recommendation as to whether the slide should be accepted or rejected for molecular sequencing. The logistic regression model can be trained on a training set of images using these derived variables. For example, the training images may be formed of pathological tissue images that have been sent and accepted for sequencing in the past, and pathological tissue images that have been rejected during regular pathology reviews. Alternatively, there may be a set threshold, for example, 20% of the nuclei on the slide may be tumors, or a minimum number of tumor cells may be required. In some examples, the model can consider the DNA ploidy of tumor cells (data from karyotyping or DNA sequence information), and by dividing the number of tumor nuclei by the average copy number of chromosomes detected in each tumor nucleus by the normally expected copy number of 2, an adjusted estimate of the available genetic material can be calculated. In some examples, the logistic regression model can be configured to have three (instead of two) possible outputs by adding an uncertainty zone between acceptance and rejection that recommends a manual review. For example, the second-to-last output of the logistic regression model is a real number, and in the last step of the model, this numerical value is thresholded at 0 to generate a binary classification. Alternatively, in some examples, the uncertainty zone is defined as a range of numerical values that includes 0. Here, values higher than this range correspond to "reject", values within the range correspond to a manual review, and values lower than this range correspond to "accept". Process 718 can be configured to calculate the size of this uncertainty zone by performing a cross-validation experiment. For example, process 718 can perform a training process that is repeated many times, but in each iteration, a different random subset of the images in the training set is used.This results in the generation of many final models that are similar but not identical, and in process 718, this variation can be used to determine the uncertainty range of the final logistic regression model. Thus, in process 718, in some examples, a binary accept / reject output can be generated, and in some examples, an accept / reject / manual review output can be generated.
[0193] Decisions are made using the recommendations from process 718. For example, the deep learning output post - processor can generate a report of the images marked as "accepted" and automatically send those images to the genomic sequencing system (112) for molecular sequencing, while the images for which "rejected" is recommended are rejected and not sent for molecular sequencing. An option for "manual review" is configured, and if recommended, the images are sent to a pathologist or a team of pathologists (118) to review the slides and determine whether to send them for molecular sequencing or reject them.
[0194] FIG. 8 shows an exemplary process 800 that can be performed by the imaging - based biomarker prediction system 102, the deep learning framework 300, or the deep learning framework 402, particularly in a deep learning framework having a single - scale configuration.
[0195] In process 802, molecular training data is received by an imaging-based biomarker prediction system. This molecular training data is for a plurality of patients and can be obtained from a gene expression dataset, such as from a source described herein. In some examples, the molecular training data includes RNA sequence data. At block 804, the molecular training data is labeled by a biomarker. One form of biomarker clustering includes labeling that can be performed by obtaining an existing label associated with a specimen, such as a tumor subtype, and associating that label with the molecular training data. Alternatively, or in addition, labeling can be performed by clustering, such as using an automated clustering algorithm. One exemplary algorithm for the case of CMS subtype biomarkers is an algorithm for identifying the CMS subtype within the molecular training data and the cluster training data according to the CMS subtype. This automated clustering can be performed, for example, within a single-class classifier module of a deep learning framework or within a slide-level label pipeline such as those within deep learning framework 300. In some examples, the molecular training data received at block 802 is RNA sequence data generated, for example, using an RNA wet lab and processed using a bioinformatics pipeline.
[0196] In various embodiments, for example, each transcriptome dataset can be generated by processing a patient or tumor organoid sample via RNA whole exome next generation sequencing (NGS) to generate RNA sequencing data, and the RNA sequence data can be processed by a bioinformatics pipeline to generate an RNA-seq expression profile for each sample. The patient sample can be a tissue sample or a blood sample containing cancer cells.
[0197] RNA can be isolated from blood samples or tissue sections using commercially available reagents such as proteinase K, TURBO DNase-I, and / or RNA Clean XP beads. The isolated RNA can be subjected to quality control protocols to determine the concentration and / or amount of RNA molecules, which may include using fluorescent dyes and a fluorescence microplate reader, a standard spectrofluorometer, or a filter fluorometer.
[0198] A cDNA library can be prepared from the isolated RNA, purified, and selected for cDNA molecule size selection using commercially available reagents such as Roche KAPA HyperBeads. Preparation of the cDNA library may include reverse transcription. In another example, a New England Biolab (NEB) kit can be used. Preparation of the cDNA library may include ligation of adapters to the cDNA molecules. For example, a UDI adapter containing a Roche SeqCap dual-end adapter, or a UMI adapter (e.g., a full-length or partial (stubby) Y-shaped adapter) can be ligated to the cDNA molecules. In this example, the adapters can function as barcodes to identify the cDNA molecules according to the samples from which they are derived and / or as barcodes to facilitate downstream bioinformatics processing and / or next-generation sequencing reactions. The sequence of nucleotides within the adapter may be unique to the sample to distinguish sequencing data obtained from different samples. The adapter may facilitate binding of the cDNA molecule to an anchor oligonucleotide molecule on the sequencer flow cell and function as a seed for the sequencing process by providing a starting point for the sequencing reaction.
[0199] The cDNA library can be amplified and purified using reagents such as Axygen MAG PCR Clean - up beads. Amplification can include polymerase chain reaction (PCR) techniques different from quantitative or reverse transcription quantitative PCR (qPCR or RT - qPCR). Next, the concentration and / or amount of cDNA molecules can be quantified using fluorescent dyes and a fluorescence microplate reader, a standard spectrofluorometer, or a filter fluorometer.
[0200] Before drying in vacuo, the cDNA libraries can be pooled and treated with reagents to reduce off - target capture, such as Human COT - 1 and / or IDT xGen Universal Blockers. Next, the pool can be resuspended in a hybridization mix, such as IDT xGen Lockdown, and probes can be added to each pool, such as IDT xGen Exome Research Panel v1.0 probes, IDT xGen Exome Research Panel v2.0 probes, other IDT probe panels, Roche probe panels, or other probes. The pool can be incubated in an incubator, PCR machine, water bath, or other temperature - controlled device to hybridize the probes. Next, the pool can be mixed with beads coated with streptavidin or another means for capturing hybridized cDNA probe molecules, particularly cDNA molecules representing exons of the human genome. In another embodiment, poly - A capture can be used. The pool can be amplified and purified again using commercially available reagents, such as the KAPA HiFi Library Amplification kit and Axygen MAG PCR Clean - up beads, respectively.
[0201] The cDNA library can be analyzed to determine the concentration or amount of cDNA molecules, for example, by using a fluorescent dye (e.g., PicoGreen pool quantification) and a fluorescence microplate reader, a standard spectrofluorometer, or a filter fluorometer. The cDNA library can also be analyzed to determine the fragment size of the cDNA molecules. This can be done via gel electrophoresis techniques, which may include using a device such as the LabChip GX Touch. The pool can be cluster amplified using a kit (e.g., the Illumina paired-end cluster kit with PhiX spike). In one example, the preparation of the cDNA library and / or the whole exome capture step can be performed in an automated system using a liquid handling robot (e.g., the SciClone NGSx).
[0202] Library amplification can be performed on a device, such as the Illumina C-Bot2, and the resulting flow cell containing the amplified target capture cDNA library can be sequenced on a next-generation sequencer, such as the Illumina HiSeq4000 or the Illumina NovaSeq6000, to a user-selected unique on-target depth, for example, up to 300x, 400x, 500x, 10,000x. The next-generation sequencer can generate FASTQ, BCL, or other files for each patient sample or each flow cell.
[0203] If two or more patient samples are processed simultaneously on the same sequencer flow cell, the reads from the multiple patient samples are initially included in the same BCL file and then split into individual FASTQ files for each patient. If the sequences of the adapters used for each patient sample are different, barcodes can serve the role of facilitating the association of each read with the correct patient sample and placement in the correct FASTQ file.
[0204] Each FASTQ file can contain paired-end or single-read reads, which can be short reads or long reads. Here, each read represents one detected nucleotide sequence within an mRNA molecule isolated from a patient sample, which is inferred by detecting the sequence of nucleotides contained in the cDNA molecules generated from the mRNA molecules isolated during library preparation using a sequencer. Each read of the FASTQ file is also associated with a quality assessment. The quality assessment may reflect the potential for errors to have affected reads associated with the sequencing procedure.
[0205] Each FASTQ file can be processed by a bioinformatics pipeline. In various embodiments, the bioinformatics pipeline can filter FASTQ data. Filtering of FASTQ data can include correction of sequencing errors and removal (trimming) of low-quality sequences or bases, adapter sequences, contaminants, chimeric reads, overrepresented sequences, biases introduced by library preparation, amplification, or capture, and other errors. Reads, individual nucleotides, or multiple nucleotides where an error can occur can be discarded based on quality assessments associated with the reads in the FASTQ file, known error rates of the sequencer, and / or comparison of each nucleotide in the read to one or more nucleotides in other reads aligned to the same position in the reference genome. Filtering can be performed in part or in whole by various software tools. The FASTQ file can be analyzed for quality control and rapid assessment of reads by sequencing data QC software such as, for example, AfterQC, Kraken, RNA-SeQC, FastQC (see Illumina's BaseSpace Labs or https: / / www.illumina.com / products / by-type / informatics-products / basespace-sequence-hub / apps / fastqc.html), or other similar software programs. For paired-end reads, the reads can be merged.
[0206] For each FASTQ file, each read within the file can be aligned to a position in the reference genome that has a sequence that most closely matches the sequence of nucleotides within the read. There are many software programs designed to align reads, such as Bowtie, the Burrows Wheeler Aligner (BWA), and programs that use the Smith-Waterman algorithm. By comparing the nucleotide sequence of each read to a portion of the nucleotide sequence of a reference genome (e.g., GRCh38, hg38, GRCh37, or other reference genomes developed by the Genome Reference Consortium), the portion of the reference genome sequence that is most likely to correspond to the read's sequence can be determined, and the reference genome can be used to direct the alignment. The alignment may take into account RNA splice sites. The alignment can generate a SAM file that stores the start and end positions of each read within the reference genome, as well as the coverage (number of reads) of each nucleotide within the reference genome. The SAM file can be converted to a BAM file, sorted, or marked for deletion of duplicate reads.
[0207] In one example, kallisto software can be used for alignment and quantification of RNA reads (see Nicolas L Bray, Harold Pimentel, Pall Melsted and Lior Pachter, "Near-optimal probabilistic RNA-seq quantification" Nature Biotechnology 34, 525-527 (2016), doi:10.1038 / nbt.3519). In another embodiment, quantification of RNA reads can be performed using another software, such as Sailfish or Salmon (see Rob Patro, Stephen M. Mount, and Carl Kingsford (2014) "Sailfish enables alignment-free isoform quantification from RNA-seq reads using lightweight algorithms" Nature Biotechnology (doi:10.1038 / nbt.2862) or Patro, R., Duggal, G., Love, M.I., Irizarry, R.A., & Kingsford, C. (2017) "Salmon provides fast and bias-aware quantification of transcript expression" Nature Methods). These RNA-seq quantification methods may not require alignment. There are numerous software packages available for normalization, quantitative analysis, and differential expression analysis of RNA-seq data.
[0208] For each gene, the raw RNA read count of a particular gene can be calculated. The raw read counts can be saved in a tabular file for each sample, where the columns represent genes and each entry represents the raw RNA read count of that gene. In one example, kallisto alignment software calculates the raw RNA read count for each read as the sum of the probabilities that the read aligns to the gene. Thus, in this example, the raw counts are not integers.
[0209] Next, for example, using full quantile normalization, raw RNA read counts can be normalized and GC content and gene length can be corrected, and the sequencing depth can be adjusted using the size factor method. In one example, normalization of RNA read counts is performed according to the methods disclosed in U.S. Patent Application No. 16 / 581,706 or PCT19 / 52801, filed on September 24, 2019, entitled "Methods of Normalizing and Correcting RNA Expression Data", which are hereby incorporated by reference in their entirety. The theoretical basis for normalization is that the copy number of each cDNA molecule in the sequencer may not reflect the distribution of mRNA molecules in the patient sample. For example, during library preparation, amplification, and capture steps, random hexamers, amplification (PCR enrichment), rRNA depletion, and various aspects of reverse transcription priming caused by probe binding and errors generated during sequencing due to the GC content, read length, gene length, and other characteristics of the sequence of each nucleic acid molecule may result in artifacts where certain portions of the mRNA molecules are over- or under-represented. Each raw RNA read count for each gene can be adjusted to eliminate or reduce over- or under-representation caused by biases or artifacts in the NGS sequencing protocol. The normalized RNA read counts can be saved in a tabular file for each sample, with columns representing genes and each entry representing the normalized RNA read count for that gene.
[0210] The set of transcriptome values can refer to either the normalized RNA read counts or the raw RNA read counts as described above.
[0211] Returning to FIG. 8, in block 804, molecular training data (e.g., such RNA sequence data) is labeled by biomarkers and clustered using an automatic clustering algorithm such as an algorithm for identifying CMS subtypes within the molecular training data and an algorithm for identifying cluster training data according to the CMS subtypes. This automatic clustering can be performed, for example, within a single-class classifier module of a deep learning framework or within a slide-level label pipeline such as those within deep learning framework 300.
[0212] In block 806, for each biomarker cluster (each corresponding to a different biomarker such as a different CMS subtype or HRD), a pathological tissue image from a related patient is obtained. These pathological tissue images can be, for example, H&E slide images with slide-level labels. In block 808, for each biomarker cluster, these labeled pathological tissue images are provided to a deep learning framework for training a biomarker classification model such as a plurality of CMS classification models for predicting different CMS subtypes. As a result, in block 810, a set of trained biomarker classifiers (classification models) is generated. Thus, blocks 802-810 represent the training process.
[0213] The prediction process starts in block 812, where a new (unlabeled or labeled) pathological tissue image such as an H&E slide image is received and provided to the single-scale biomarker classifier generated by block 810. In block 814, the biomarker classification on the received pathological tissue image is predicted to be determined by one or more biomarker classification models such as one or more CMS subtypes or HRD.
[0214] Similar to block 610, at block 814, a new histopathological image can be received from a physical clinical record system or a primary care system and applied to a trained deep learning framework that determines biomarker predictions with its tissue classification model and / or biomarker classification model. The prediction score can be determined, for example, for the entire histopathological image.
[0215] Further, as shown in process 900 of FIG. 9, similar to process 600, after prediction, the predicted biomarker classification from block 814 can be received at block 902. At block 904, a histopathological image, and thus a clinical report for the patient, can be generated, which includes the status of the predicted biomarker. At block 906, an overlay map indicating the predicted biomarker status can be generated, which is provided to a clinician for display or to a pathologist to determine a preferred immunotherapy corresponding to the predicted biomarker.
[0216] FIGS. 10A and 10B show examples of digital overlay maps created by, for example, the overlay map generator 324 of system 300. These overlay maps can be generated as static digital reports for display to a clinician or as dynamic reports that enable user interaction via a graphical user interface (GUI). FIG. 10A shows a tissue class overlay map generated by the overlay map generator 324. FIG. 10B shows a cell outline overlay map generated by the overlay map generator 324.
[0217] In one example, the overlay map generator 324 can display a digital overlay as a transparent or opaque layer covering the pathological tissue image, aligned such that the image position shown on the overlay and the pathological tissue image are at the same position on the display. The transparency of the overlay map can be various. The transparency can be adjustable by the user in the dynamic reporting mode of the overlay map generator 324. The overlay map generator 326 can report the percentage of labeled tiles associated with each tissue class label, the ratio of the number of tiles classified into each tissue class, the total area of all grid tiles classified into a single tissue class, and the ratio of the area of tiles classified into each tissue class. The overlay map can show various tissue classifications and can be displayed as a heat map having various pixel intensity levels corresponding to various biomarker status levels. For example, in the case of TIL, pixels with high intensity in tissue regions with a high predicted TIL status (% is high) and pixels with low intensity in tissue regions with a low predicted TIL status (% is low) are shown.
[0218] In one example, the deep learning output post - processor 308 can also report the total number or percentage of cells in a region defined by any one of the user, the entire slide, a single grid tile, all grid tiles classified into each tissue class, or cells classified as immune cells. The controller 308 can also report the number of cells classified as lymphocyte cells located within a region classified as a tumor or any other tissue class.
[0219] In one example, the digital overlays and reports generated by controller 308 can be used to assist in localizing or diagnosing a region of interest of a subject, including invasive tumors having tumor cells that protrude into non-tumor tissue regions surrounding the tumor, by enabling a medical professional to more accurately estimate tumor purity. They can also assist a medical professional in prescribing treatment. For example, the number of lymphocytes in a region classified as a tumor can be used to predict whether immunotherapy will be successful in treating a patient's cancer.
[0220] In one example, the digital overlays and reports generated by controller 308 can be used to determine whether a slide sample has high-quality tissue sufficient for successful genetic sequence analysis of the tissue, e.g., whether to implement an accept / reject / manual determination as discussed in process 700. Genetic sequence analysis of the tissue on the slide may succeed if the slide contains a certain amount of tissue and / or if there is a tumor purity value above user-defined thresholds for tissue amount and tumor purity. Controller 308 can use process 700 to label a slide as accepted or rejected for sequence analysis according to the amount of tissue present on the slide and the tumor purity of the tissue on the slide. Controller 308 can also, in accordance with process 700, label a slide as uncertain using user-defined tissue amount thresholds and user-defined uncertainty ranges obtained from a user interacting with the digital overlays and reports from generator 324.
[0221] In one example, for instance, using the biomarker metric processing module 326, the controller 308 implementing the process 700 calculates the amount of tissue on the slide by measuring the total area covered by the tissue in the pathological tissue image or by counting the number of cells on the slide. The number of cells on the slide can be determined by the number of cell nuclei visible on the slide. In one example, the controller 308 calculates the proportion of tissue that is cancerous by dividing the number of cell nuclei within the grid area labeled as tumor by the total number of cell nuclei on the slide. The controller 308 can exclude the outer edges of cell nuclei or cells that belong to cells located in the tumor area but are characterized as lymphocytes. The proportion of tissue that is cancerous is known as the tumor purity of the sample. Next, the controller 308 compares the tumor purity with a minimum tumor purity threshold selected by the user and compares the number of cells in the digital image with a minimum cell threshold selected by the user (input by the user interacting with the overlay map generator 324). If both thresholds are exceeded, the tissue slide depicted in the image is approved for molecular testing, including gene sequence analysis. In one example, the minimum tumor purity threshold selected by the user is 0.20, i.e., 20%. However, any number of tumor purity thresholds can be selected, including 1%, 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, or more.
[0222] In another example, a composite tissue amount score is provided for the image by multiplying the total area covered by the tissue detected on the slide by the controller 308 by a first multiplier value, multiplying the number of cells counted on the slide by a second multiplier value, and summing the products of these multiplications.
[0223] In one example, the controller 308 can calculate whether the tumor and the labeled grid regions are spatially integrated or dispersed among non-tumor grid regions. If the controller 308 determines that the tumor region is spatially integrated, the overlay map generator 324 can generate a digital overlay of the recommended cut boundary, which separates the image regions classified as tumor and non-tumor, or the image regions proximal to the regions classified as tumor within the regions classified as non-tumor. This recommended cut boundary can serve as a guide to assist the technician in dissecting the slide to separate the maximum amount of tumor or non-tumor tissue from the slide, particularly for molecular examinations including gene sequence analysis.
[0224] In one example, the controller 308 can include a clustering algorithm, which calculates and reports information regarding the spacing and density of type-classified cells, tissue-classified tiles, or visually detectable features on the slide. The spacing information includes the distribution patterns and heat maps of lymphocytes, immune cells, tumor cells, or other cells. These patterns may include clustering, dispersion, high density, and absence. This information helps determine whether immune cells and tumor cells cluster together and what percentage of the cluster regions overlap, which can facilitate predicting the patient's response to immune infiltration and immunotherapy.
[0225] The controller 308 can also calculate and report the average tumor cell circularity, average tumor cell perimeter length, and average tumor nuclear density.
[0226] The spacing information also includes the mixing level of tumor cells and immune cells. The clustering algorithm can calculate the probability that two adjacent cells on a given slide or within a region of the slide will be either two tumor cells, two immune cells, or one tumor cell and one immune cell.
[0227] The clustering algorithm can also measure the thickness of some stromal patterns located around the areas classified as tumors. The thickness of the stroma surrounding this tumor area can be a predictor of the patient's response to treatment.
[0228] In one example, the controller 308 can also calculate and report statistics including mean, standard deviation, sum, etc. for red, green, blue (RGB) values, luminance, hue, saturation, grayscale, and deconvolution of staining within each grid tile aggregated from a single slide image or multiple slide images. Deconvolution includes the removal of visual signals created by several individual stains or combinations of stains, including hematoxylin, eosin, or IHC staining.
[0229] The controller 308 can also incorporate known mathematical formulas from the fields of physics and image analysis to calculate visually detectable basic features of each grid tile. By combining visually detectable basic features including lines, alternating brightness patterns, and shapes that can draw contours, visually detectable complex features can be created, including cell size, cell circularity, cell shape, and staining patterns called texture features.
[0230] In other examples, the digital overlays, reports, statistics, and estimates generated by the overlay map generator 324 can be useful for predicting patient survival, the patient's response to a particular cancer treatment, the PD-L1 status of a tumor or immune cluster, microsatellite instability (MSI), tumor mutation burden (TMB), and the origin of the tumor if the origin of the tumor is unknown or the tumor is metastatic. The biomarker metric processing module 326 can also calculate quantitative measurements of the predicted patient survival, the patient's response to a particular cancer treatment, the PD-L1 status of a tumor or immune cluster, microsatellite instability (MSI), and tumor mutation burden (TMB).
[0231] In one example, the controller 308 can calculate the relative density of each type of immune cell over the entire slide in an area designated as a tumor or another tissue class. Immune tissue classes include lymphocytes, cytotoxic T cells, B cells, NK cells, macrophages, and the like.
[0232] In one example, the act of scanning or otherwise digitally capturing a pathological tissue slide automatically triggers the deep learning framework 300 to analyze the digital image of the pathological tissue slide.
[0233] In one example, the overlay map generator 324 enables a user to edit the extracellular edge or boundary between two tissue classes on a tissue class overlay map or an extracellular edge overlay map and saves the modified map as a new overlay.
[0234] FIG. 11 shows a process 1100 for preparing a digital image of a pathological tissue slide for tissue classification, biomarker detection, and mapping analysis, which can be implemented using the system 300. The process 1100 can be executed for each received image for analysis and biomarker prediction. In some examples, the process 1100 can be executed, in whole or in part, on the initially received training images. Each of the processes described in FIG. 9 can be executed by the preprocessing controller 302, where any one or more of the processes can be executed by the normalization module 310 and / or the tissue detector 314.
[0235] In an example such as when training a classifier model, each digital image file received at 1102 by the preprocessing controller 302 includes multiple versions of the same image content, and each version has a different resolution. The file stores these copies in a stack layer and arranges them by resolution such that the highest resolution image containing the maximum number of bytes is at the bottom layer. This is known as a pyramid structure. In one example, the highest resolution image is the highest resolution achievable by the scanner or camera that created the digital image file.
[0236] In one example, each digital image file also includes metadata indicating the resolution of each layer. The preprocessing controller 302 can, in process 1104, detect the resolution of each layer in this metadata and compare it with the resolution criteria selected by the user to select the layer with the optimal resolution for analysis. In one example, the optimal resolution is 1 pixel per micron (downsampled by 4).
[0237] In one example, the preprocessing controller 302 receives a tagged image file format (TIFF) file with a bottom layer resolution of 4 pixels per micron. This resolution of 4 pixels per micron corresponds to the resolution achieved by a microscope objective lens with a magnification of "40x". In one example, the area on the slide where tissue may be present is a maximum size of 100,000 × 100,000 pixels.
[0238] In one example, the TIFF file has approximately 10 layers, and the resolution of each layer is half the resolution of the layer below it. If the resolution of the high-resolution layer is 4 pixels per micron, the resolution of the layer above it will be 2 pixels per micron. The area represented by 1 pixel in the upper layer is the size of the area represented by 4 pixels in the lower layer, that is, the length of each side of the area represented by 1 upper layer pixel is twice the length of each side of the area represented by 1 lower layer pixel.
[0239] Each layer can be downsampled by a factor of two with respect to the layer below it, as executed in process 1106. Downsampling is a method by which a new version of the original image can be created with a lower resolution value than the original image. There are many methods known in the art for downsampling, including nearest neighbor, bilinear, Hermite, Bell, Mitchell, bicubic, and Lanczos resampling.
[0240] In one example, downsampling by a factor of two means that out of four pixels in a square of the high-resolution layer, the red, green, and blue (RGB) values of three of them are replaced by the RGB value of the fourth pixel, creating a new large pixel in the upper layer that occupies the same space as the averaged four pixels.
[0241] In one example, the digital image file does not include a layer or image at the optimal resolution. In this case, the preprocessing controller 302 can receive an image from a file having a resolution higher than the optimal resolution in process 1106 and downsample the image at a ratio to achieve the optimal resolution.
[0242] In one example, the optimal resolution is two pixels per micron, i.e., a magnification of "20x", but the bottom layer of the TIFF file is four pixels per micron, and each layer is downsampled by a factor of four compared to the layer below it. In this case, the TIFF file has one layer at a magnification of 40x and the next layer at a magnification of 10x, but no layer at a magnification of 20x. In this example, the preprocessing controller 302 reads the metadata, compares the resolution of each layer to the optimal resolution, and cannot find a layer at the optimal resolution. Instead, the preprocessing controller 302 searches for the layer at a magnification of 40x and then downsamples the image of that layer at a downsampling ratio of two to create an image with the optimal resolution of 20x magnification.
[0243] Also, in process 1106, after the preprocessing controller 302 acquires an image at an optimal resolution, it identifies all parts of the image that depict the tumor sample tissue and digitally removes fragments, pen marks, and other non-tissue objects.
[0244] In one example, also in process 1106, the preprocessing controller 302 differentiates between the tissue and non-tissue regions of the image and uses Gaussian blur removal to edit the pixels having non-tissue objects. In one example, any control tissue on the slide that is not part of the tumor sample tissue can be detected by a tissue detector and labeled as control tissue or manually labeled by a human analyst as control tissue to be excluded from downstream tile grid projection.
[0245] Non-tissue objects include artifacts, markings, and fragments within the image. Fragments include keratin, highly compressed or fragmented tissue that cannot be visually analyzed, and objects not collected in the sample.
[0246] In one example, also in process 1106, the slide image contains marker ink or other writing, and the controller 302 detects and digitally deletes it. The marker ink or other writing may be transparent on the tissue, i.e., the tissue on the slide may be visible through the ink. Since the ink of each marking is of one color, the ink causes a consistent shift in the RGB values of the pixels containing the tissue stained under the ink as compared to the pixels containing the tissue stained without ink.
[0247] In one example, also in process 1106, the controller 302 identifies the portion of the slide image that contains ink by detecting portions having RGB values different from the RGB values of the remaining portion of the slide image, and the difference in RGB values from the two portions is consistent. Next, the tissue detector can digitally remove the ink by subtracting the difference between the RGB values of the pixels in the ink portion and the pixels in the non-ink portion from the RGB values of the pixels in the ink portion.
[0248] In one example, also in process 1106, the controller 302 excludes pixels within the image that have little local variation. These pixels represent artifacts, markings, or blurred regions caused by out-of-focus tissue slices, air bubbles trapped between two glass layers of the slide, or pen marks on the slide.
[0249] In one example, in process 1106 as well, the controller 302 removes these pixels by converting the image to a grayscale image, passes the grayscale image through a Gaussian blur filter, and mathematically adjusts the original grayscale value of each pixel to a blurred grayscale value to create a blurred image. Other filters can be used to blur the image. Next, for each pixel, the controller 302 subtracts the blurred grayscale value from the original grayscale value to create a differential grayscale value. In one example, if the differential grayscale value of a pixel is less than a user-defined threshold, it may indicate that the blur filter did not significantly change the original grayscale value and that the pixel was in the blurred region of the original image. By comparing the differential grayscale value to the threshold, a binary mask can be created that indicates where there are blurred regions that can be designated as non-tissue regions. The mask can be a copy of the image, with the color, RGB value, or other value within the pixel adjusted to indicate the presence or absence of a particular type of object and to indicate the location of all objects of that type. For example, the binary mask can be generated by setting the binary value of each pixel to 0 if the pixel has a differential grayscale value less than the user-defined blur threshold and setting the binary value of each pixel to 1 if the pixel has a differential grayscale value greater than or equal to the user-defined blur threshold. The regions of the binary mask where the pixel binary value is 0 indicate the blurred areas of the original image that may be designated as non-tissue.
[0250] The controller 302 can also mute or remove extreme brightness or darkness in the image in process 1108. In one example, the controller 302 converts the input image into a grayscale image, and each pixel receives a numerical value according to the brightness of the pixel. In one example, the range of grayscale values is 0 to 255, where 0 represents black and 255 represents white. For pixels with grayscale values exceeding the brightness threshold, the tissue detector replaces the grayscale values of those pixels with the brightness threshold. In the case of pixels whose grayscale values are below the darkness threshold, the tissue detector replaces the grayscale values of those pixels with the darkness threshold. In one example, the brightness threshold is about 210. In one example, the darkness threshold is about 45. The tissue detector stores the image with the new grayscale values in a data file.
[0251] In one example, the controller 302 analyzes the modified image for artifacts, debris, or markings remaining after the initial analysis in process 1110. The tissue detector scans the image and classifies the remaining pixel groups with specific colors, sizes, or smoothness as non-tissue.
[0252] In one example, the slide has an H&E stain, and most of the tissue in the pathological tissue image has a pink stain. In this example, the controller 302 classifies all objects without a pink or red hue as non-tissue as determined by the RGB values of the pixels representing the objects. The tissue detector 314 can interpret any color or lack of any color within a pixel to indicate the presence or absence of tissue within that pixel.
[0253] In one example, the controller 302 detects the outer shape of each object in the image to measure the size and smoothness of each object. Very dark pixels may be debris, and very bright pixels may be the background, both of which are unstructured objects. Thus, the controller 302 can detect the outer shape of each object by converting the image to grayscale, comparing the grayscale value of each pixel with a range of values that the user has determined is not too bright or too dark, determining whether the grayscale value is within the range, and generating a binary image in which each pixel is assigned one of two numerical values.
[0254] For example, to threshold the image, the controller 302 can compare the grayscale value of each pixel with a range of user-defined values, replace each grayscale value outside the user-defined range with the value 0, and replace each grayscale value within the user-defined range with the value 1. Next, the controller 302 draws the outer shape of each object as the outer edge of each group of adjacent pixels having the value 1. The closed outer shape indicates the presence of an object, and the controller 302 measures the size of the object by measuring the area within the outer shape of each object.
[0255] In one example, the likelihood that a tissue object on the slide contacts the outer edge of the slide is low, and the controller 302 classifies any object that contacts the edge of the slide as non-tissue.
[0256] In one example, after measuring the size of each object, the controller 302 ranks the sizes of all the objects and designates the maximum value as the size of the largest object. The controller 302 divides the size of each object by the size of the largest object and compares the resulting size quotient with a user-defined size threshold. If the size quotient of an object is smaller than the user-defined size threshold, the controller 302 designates that object as non-tissue. In one example, the user-defined size threshold is 0.1.
[0257] Before measuring the size of each object, in process 1106, the controller 302 can first downsample the input image to reduce the possibility that a part of the tissue object is designated as non-tissue. For example, a single tissue object may appear as a first tissue object part surrounded by one or more additional tissue object parts having a smaller size. After thresholding, the additional tissue object parts may have a size quotient smaller than the user-defined size threshold and may be erroneously designated as non-tissue. By downsampling before thresholding, a small group of adjacent pixels with a value of 1 surrounded by pixels with a value of 0 in the original image is included in a larger group of proximal pixels with a value of 1. The opposite may also be true if a small group of adjacent pixels with a value of 0 surrounded by pixels with a value of 1 in the original image is included in a larger group of proximal pixels with a value of 0.
[0258] In one example, the controller 302 downsamples an image with a magnification of 40 times at a ratio of 16 times, so the magnification of the resulting downsampled image is 40 / 16 times, and each pixel of the downsampled image represents 16 pixels of the original image.
[0259] In one example, in process 1110, the controller 302 detects the boundaries of each object on the slide as clusters of pixels surrounded by pixels having a binary value or RGB value not equal to zero and having an RGB value equal to zero that indicates the boundary of the object. If the pixels forming the boundary are relatively on a straight line, the controller 302 classifies the object as unstructured. For example, the controller 302 uses a closed polygon to draw the contour of a shape. If the number of vertices of the polygon is less than a user-defined minimum vertex threshold, the polygon is considered to be a simple inorganic shape that is too smooth and is marked as unstructured. Next, in process 1112, the controller 302 applies a tiling process to the normalized image.
[0260] Figures 12A - 12C show an exemplary architecture 1200 that can be used for the classification model of module 306. For example, the same architecture 1200 can be used for each of the tissue segmentation model 322 and the tissue classification model 320, both of which are implemented using the FCN configuration or any neural network herein. The tissue classifier module 306 includes a tissue classification algorithm (see Figures 12A - 12C) that assigns tissue class labels to the images represented by each received tile (an exemplary tile 1302 is labeled in the first portion 1304 of the pathological tissue image 1300 shown in Figure 13). In one example, the overlay map generator 324 can report the assigned tissue class labels associated with each small square tile by displaying a grid-based digital overlay map in which each tissue class is represented by a unique color (see Figure 12A).
[0261] If the tile size is small, the time required for the tissue classifier module 306 to analyze the input image may increase. Alternatively, if the tile size is large, it is more likely that two or more tissue classes will be included in the tile, and it may become difficult to assign a single tissue class label to the tile. In this case, instead of calculating that the probability of one of the tissue class labels describing the image in a small square tile is high compared to other tissue class labels, the architecture 1200 can calculate that the probabilities of two or more tissue class labels being accurately assigned to a single small square tile are equal.
[0262] In one example, each side of each small square tile is about 32 microns in length, and about 5 - 10 cells fit into each small square tile. This small tile size allows the tissue classifier module 306 to create a more spatially accurate boundary when determining the boundary between two adjacent small square tile regions representing two different tissue classes. In one example, each side of the small square tile can be as short as 1 micron.
[0263] In one example, the size of each tile can be set by the user to contain a specific number of pixels. In this example, the length of each side of the tile measured in microns will be determined by the resolution of the input image. At different resolutions, the length of the tile side in microns will be different, and the number of cells in each tile may be different.
[0264] The architecture 1200 recognizes various pixel data patterns of the portion of the digital image located within or near each small square tile and assigns a tissue class label to each small square tile based on those detected pixel data patterns. In one example, a medium-sized square tile centered on the small square tile includes an area of the slide image close enough to the small square tile to contribute to the label assignment of that small square tile.
[0265] In one example, each side of the medium-sized square tile is about 466 microns in length, and each medium-sized square tile contains about 225 (15×15) small square tiles. In one example, this medium tile size increases the likelihood that the features of the tissue structure can fit within a single medium tile, providing context to the algorithm when labeling the central small square tile. The features of the tissue structure may include glands, ducts, blood vessels, immune clusters, and the like.
[0266] In one example, this medium tile size is selected such that it can cancel out the shrinkage that occurs during convolution.
[0267] During convolution by architecture 1200, the input image matrix is multiplied by a filter matrix to create a result matrix. Shrinkage refers to the case where the result matrix is smaller than the input image matrix. The dimensions of the filter matrix of the convolutional layer affect the number of rows and columns lost due to shrinkage. The total number of matrix entries lost due to shrinkage by processing an image through a specific CNN can be calculated according to the number of convolutional layers of the CNN and the dimensions of the filter matrix of each convolutional layer (see FIGS. 12A - 12C).
[0268] In the example shown in FIG. 12B, the combined convolutional layers lose a total of 217 rows or columns of the matrix from the top, bottom, and two side edges of the matrix. Thus, the medium square tile is set to be equal to the small square tile with 217 pixels added to both sides of the small square tile.
[0269] In one example, two adjacent small square tiles share a side and each is at the center of a medium-sized square tile. Two medium-sized square tiles overlap. Of the 466*466 small pixels in each medium-sized square tile, the two medium-sized square tiles share all but 32*466 pixels. In one example, each convolutional layer of the algorithm (see FIGS. 12A and 12B) analyzes both medium-sized square regions simultaneously such that the algorithm generates a vector of two values (one for each of the two small square tiles).
[0270] The vector of values includes probability values for each tissue class label, which indicates the likelihood that the small square tile represents that tissue class. The vectors of values can be arranged in a matrix form to form a three-dimensional probability data array. The location of each vector within the three-dimensional probability data array relative to the other vectors corresponds to the location of the associated small square tile relative to the other small square tiles included in the algorithm analysis.
[0271] In this example, 434×434 (188, 356) of the 466×466 (217, 156) pixels of each medium-sized square tile are common to both medium-sized square tiles. By analyzing both medium-sized square tiles simultaneously, the algorithm increases efficiency.
[0272] In one example, architecture 1200 can further increase efficiency by analyzing a large tile formed by a plurality of overlapping medium-sized square tiles, each of which includes a number of small square tiles surrounding one central small square tile that receives a label of a tissue class. In this example, the algorithm generates one data structure in the form of a three-dimensional probability data array that includes one vector of probabilities for each small square tile, and the location of the vectors within the three-dimensional array corresponds to the location of the small tiles within the large tile.
[0273] Architecture 1200 stores this three-dimensional probability data array, for example, within the tissue classifier module 306, and the overlay map generator 324 converts the tissue class label probabilities for each small square tile into a tissue class overlay map. In one example, the overlay map generator 324 can compare the probabilities stored in each vector to determine the maximum probability value associated with each small square tile. The tissue class label associated with that maximum value can be assigned to that small square tile, and only the assigned labels will be displayed in the tissue class overlay map.
[0274] In one example, the matrices generated by each layer of architecture 1200 for the large square tiles are stored in graphics processing unit (GPU) memory. The possible maximum size of the large square tiles can be determined by the capacity of the GPU memory and the amount of GPU memory required for each entry in the three-dimensional probability data array. In one example, the GPU memory capacity is 250 MB, and 4 bytes of GPU memory are required for each entry in the matrix. This allows for a large tile size of 4,530 pixels × 4,530 pixels, calculated as follows: 4 bytes / entry * 4530 * 4530 * 3 entries per large tile = 246 (approx. 250) MB of GPU memory required per large square tile. In another example, 8 bytes of GPU memory are required for each entry in the matrix. In this example, a 16 GB GPU can process 32 large tiles simultaneously, and the size of each large tile has dimensions of 4,530 pixels × 4,530 pixels, calculated as follows: 32 large tiles * 8 bytes / entry * 4530 * 4530 * 3 entries per large tile = 14.7 (approx. 16) GB of GPU memory required.
[0275] In one example, each entry in the three-dimensional probability data array is a data entry in single-precision floating-point format (float32).
[0276] In one example, there are 16,384 (128^2) non - overlapping small square tiles that form a large square tile. Each small square tile is at the center of a medium - sized square tile with sides approximately 466 pixels in length. The small square tiles form the central region of a large square tile with sides approximately 4,096 pixels in length. All the medium - sized square tiles overlap and create a border approximately 217 pixels wide around all four sides of the central region. Including the border, each large square tile has sides approximately 4,530 pixels in length.
[0277] In this example, this large square tile size enables parallel computing and reduces the percentage of redundant calculations by 99%. This can be calculated as follows. First, a pixel inside the large square tile (any pixel at least 434 pixels from the edge of the large square tile) is selected, and centered on this model pixel, a region the size of a medium - sized square tile (466 pixels at the edge) is created. Then, for the small square tile at the center of this constructed region, the model pixel is included within the corresponding medium - sized square tile of that small square tile. There are (466 / 32)^2 = approximately 217 small square tiles like this within a large square tile. For pixels not inside the large square tile, the number of small square tiles that meet this condition is reduced. As the distance between the selected small square tile and the edge of the large square tile decreases, the number decreases linearly, and then as the distance between the selected small square tile and the corner decreases, a small number of pixels (about 0.005%) contribute only to the classification of a single small square tile. Performing the classification with a single large square tile means that the calculation for each pixel is done only once, not once per small square tile. Thus, the redundancy is reduced by approximately a factor of 217. In one example, since there may be several large square tiles on a slide, each potentially overlapping slightly with adjacent tiles, the redundancy is not completely eliminated.
[0278] The upper limit of the redundant calculation percentage can be set (a slight deviation from this upper limit varies depending on the number of large square tiles required to cover the organization and the relative arrangement of these tiles). The percentage of redundancy is 1 - 1 / r, where r is the redundancy ratio, and r can be calculated as (T / N + 1)(sqrt(N)*E + 434)^2 / (sqrt(T)*E + 434)^2, where T is the total number of small square tiles on the slide, N is the number of small square tiles per large square tile, and E is the edge size of the small square tile.
[0279] Figure 12A shows a layer of an example layer structure of architecture 1200. Figure 12B shows an example of the output sizes of different layers of architecture 1200 and the resulting sub-layers, showing the FCN configuration at tile resolution. As shown, the tile resolution FCN configuration included in the tissue classifier module 306 has additional layers of 1×1 convolution with skip connections, downsamples 8-fold with skip connections, uses a confidence map layer, replaces the average pooling layer with a concatenation layer, and replaces the fully connected FCN layer with a 1×1 convolution and a softmax layer. The additional layers convert the classification task into a classification segmentation task. This means that instead of receiving and classifying the entire image as one tissue class label, the additional layers enable the tile resolution FCN to classify each small tile within the user-defined grid as a tissue class.
[0280] These additional and replacement layers convert the CNN into a tile resolution FCN without requiring upsampling implemented in the layers after the conventional pixel resolution FCN. Upsampling is a method that can create a new version of the original image with a higher resolution value than the original image. However, upsampling is a time-consuming and computationally intensive process, which can be avoided in this architecture.
[0281] There are many methods known in the art for upsampling, including nearest neighbor, bilinear, Hermite, Bell, Mitchell, bicubic, and Lanczos resampling. In one example, 2x upsampling means that a pixel with red, green, blue (RGB) values is split into four pixels, and the RGB values of three new pixels can be selected to match the RGB values of the original pixel. In another example, the RGB values of the three new pixels can be selected as the average of the RGB values from the original pixel and the pixels adjacent to the adjacent pixels.
[0282] Since the RGB values of the new pixels may not accurately reflect the visible tissue of the original slide captured by the digital slide image, upsampling can introduce errors into the final image overlay map generated by the overlay map generator 224.
[0283] In one example, instead of labeling individual pixels, the tile resolution FCN is programmed to analyze large square tiles made up of small square tiles, and generates a 3D array of values each representing the probability that one tissue class classification label matches the tissue class depicted in each small tile. The convolutional layers known in the art perform a multiplication of at least one filter matrix on at least one input image matrix. In the first convolution later, the input image matrix has the values of all the pixels of the large square tile input image and represents the visual data of those pixels (e.g., values from 0 to 255 for each channel of RGB).
[0284] The filter matrix can have dimensions selected by the user and can include weight values selected by the user or determined by backpropagation during the training of the CNN model. In one example, in the first convolutional layer, the dimensions of the filter matrix are 7×7 and there are 64 filters. The filter matrix may represent a visual pattern that can distinguish one tissue class from another.
[0285] In an example where RGB values are input into the input image matrix, the input image matrix and the filter matrix are three-dimensional (see FIG. 12C). Each input image matrix is multiplied by each filter matrix to generate a result matrix. All the result matrices generated by the filters of one convolutional layer can be stacked to create a three-dimensional result matrix with dimensions such as rows, columns, and depth. The depth, which is the last dimension of the 3D result matrix, has a depth equal to the number of filter matrices. The result matrix from one convolutional layer becomes the input image matrix for the next convolutional layer.
[0286] Returning to FIG. 12A, the title of the convolutional layer containing “ / n” (where n is a numerical value) indicates that there is downsampling (known as pooling) of the result matrix generated by that layer. n indicates the factor by which downsampling occurs. A 2-fold downsampling means that the downsampled result matrix with half the number of rows and half the number of columns of the original result matrix is created by replacing the square of four values of the result matrix with one of those values, or a statistic calculated from those values. For example, the minimum value, maximum value, or average value of the values may be replaced with the original values.
[0287] Architecture 1200 also adds skip connections (shown in FIG. 12A as the black line with an arrow connecting the blue convolutional layer directly to the connection layer). The left skip connection includes an 8-fold downsampling, and the right skip connection includes two convolutional layers that multiply the input image matrix by filter matrices each having a 1×1 dimension. Since the filter matrices of these layers are 1×1 in dimension, only the corresponding probability vectors of the individual small square tiles of the result matrix created by the purple convolutional layer contribute. These result matrices represent a small field of view.
[0288] In all other convolutional layers, due to the large dimensions of the filter matrix, the pixels of each medium-sized square tile that contains a small square tile at the center of the medium-sized square tile can contribute to the probability vector of the resulting matrix corresponding to that small square tile. These resulting matrices can influence the probability that the context pixel data pattern surrounding the small square tile applies the label of each tissue class to the small square tile. These resulting matrices represent a large field of view.
[0289] With the 1×1 convolutional layer of the skip connection, the algorithm can consider the pixel data pattern of the small central square tile to be more or less important than the remaining pixel data patterns of the surrounding medium-sized square tiles. This is reflected by the weights by which the trained model is multiplied by the final resulting matrix from the convolutional layer of the medium-sized tiles between the concatenation layers (shown in the central column of Figure 10A) compared to the weights by which the trained model is multiplied by the final resulting matrix from the skip connection layer (shown on the right side of Figure 12A).
[0290] The downsampling skip connection shown on the left side of Figure 12A creates a resulting matrix with a depth of 64. The 3×3 convolutional layer with 512 filter matrices creates a resulting matrix with a depth of 512. The 1×1 convolutional layer with 64 filter matrices creates a resulting matrix with a depth of 64. All three of these resulting matrices have the same number of rows and the same number of columns. The concatenation layer concatenates these three resulting matrices to form a final resulting matrix with the same number of rows, the same number of columns, and a depth of 64 + 512 + 64 (640) as the three concatenated matrices. This final resulting matrix combines the focus on the magnitude of the view matrices.
[0291] The final resulting matrix can be flattened into two dimensions by multiplying a coefficient to all entries and summing the products along each depth. Each element can be selected either by the user or by backpropagation during model training. Flattening does not change the number of rows and columns of the final resulting matrix, but the depth is changed to 1.
[0292] The 1×1 convolutional layer receives the final result matrix and filters it with one or more filter matrices. The 1×1 convolutional layer can include one filter matrix associated with each tissue class label of the trained algorithm. This convolutional layer generates a 3D result matrix having a depth equal to the number of tissue class labels. Each depth corresponds to one filter matrix, and along the depth of the result matrix, there may be probability vectors for each small square tile. This 3D result matrix is a three-dimensional probability data array, and the 1×1 convolutional layer stores this 3D probability data array.
[0293] The softmax layer can create a two-dimensional probability matrix from the 3D probability data array by comparing all the values of each probability vector, selecting the tissue class associated with the maximum value, and assigning that tissue class to the small square tile associated with that probability vector.
[0294] Next, the stored three-dimensional probability data array or 2D probability matrix can be converted into a tissue class overlay map in the final confidence map layer of FIG. 10A to efficiently assign tissue class labels to each tile.
[0295] In one example, to cancel shrinkage, rows and columns are added to all four outer edges of the input image matrix, and each value entry in the added rows and columns is zero. These rows and columns are called padding. In this case, the training data input matrix will have the same number of added rows and columns with value entries equal to zero. Differences in the number of padding rows or columns of the training data input matrix will result in filter matrix values that do not cause the tissue class locator 216 to mislabel the input image.
[0296] In the FCN shown in FIG. 12A, since there are gray and blue layers, 217 total outer rows or columns on each side of the input image matrix are lost due to contraction before the skip connection. Only the pixels in the small square tiles have vectors corresponding to the resulting matrix created after the green layer.
[0297] In one example, each medium-sized square tile is not padded by adding rows and columns with zero-valued entries around the input image matrix corresponding to each medium-sized square tile. This is because the zeros would replace the image data values from adjacent medium-sized square tiles that the tissue classifier 216 needs to analyze. In this case, the training data input matrix is also not padded.
[0298] FIG. 12C visualizes each depth of an exemplary three-dimensional input image matrix being convolved by two exemplary three-dimensional filter matrices.
[0299] In an example where the RGB channels of each medium-sized square tile are included in the input image matrix, the input image matrix and the filter matrix become three-dimensional. In one of the three dimensions, the input image matrix and each filter matrix have three depths for the red channel, the green channel, and the blue channel.
[0300] The red channel (first depth) 1202 of the input image matrix is multiplied at the corresponding first depth of the first filter matrix. The green channel (second depth) 1204 is multiplied in a similar manner, and the blue channel (third depth) 1206 is also multiplied in a similar manner. Next, the red, green, and blue product matrices are summed to create the first depth of the three-dimensional result matrix. This is repeated for each filter matrix to create additional depths of the three-dimensional result matrix corresponding to each filter.
[0301] A wide variety of training sets can be used to train the CNN or FCN included in the tissue classifier module 306.
[0302] In one example, the training set may include JPEG images of medium-sized square tiles, each having a tissue class label assigned to its central small square tile, obtained from at least 50 digital images of a pathological tissue slide at a resolution of approximately 1 pixel per micron. In one example, a human analyst delineates the contours of all relevant tissue classes and labels (annotates the various tissue classes) or labels each small square tile of each pathological tissue slide as non-tissue or a specific type of cell. Tissue classes may include tumor, stroma, normal, immune cluster, necrosis, hyperplasia / dysplasia, and erythrocytes. In one example, each side of each central small square tile is approximately 32 pixels in length.
[0303] In one example, the training set images are converted into an input training image matrix and processed by a tissue classifier module 306 such that a tissue class label is assigned to each tile image of the training image. If the tissue classifier module 306 does not accurately label the validation set of the training images to match the corresponding annotations added by a human analyst, the weights of each layer of the deep learning network can be automatically adjusted by stochastic gradient descent by backpropagation until the tissue classifier module 306 accurately labels most of the validation set of the training images.
[0304] In one example, the training dataset has multiple classes, where each class represents a tissue class. The training set can generate a unique model using specific hyperparameters (number of epochs, learning rate, etc.) that can recognize the content of digital slide images and classify them into various classes. Tissue classes may include tumor, stroma, immune cluster, normal epithelium, necrosis, hyperplasia / dysplasia, and erythrocytes. In one example, if there is a sufficient training set for each tissue class, the model can classify an unlimited number of tissue classes.
[0305] In one example, the training set images are converted into grayscale masks for annotation, where different values (0 - 255) of the mask image represent different classes.
[0306] Since each histopathological image can exhibit significant variations in visual features, including the appearance of tumors, the training set may contain highly different digital slide images to more appropriately train a model for a wide variety of slides that may be analyzed. Images within the training data may be subjected to data augmentation (including rotation, scaling, color jitter, etc.) before being used for model training.
[0307] The training set can also be specific to the type of cancer. In this case, all histopathological slides from which digital images are generated for a particular training set contain tumor samples from the same type of cancer. Types of cancer can include breast, colorectal, lung, pancreas, liver, stomach, skin, etc. Each training set can create a unique model specific to the type of cancer. Each type of cancer can also be divided into subtypes of cancer known in the art.
[0308] In one example, the training set can be derived from a pair of pathological tissue slides. The pair of pathological tissue slides includes two pathological tissue slides, each having one slice of the tissue, and the two slices of the tissue were arranged substantially proximally / almost adjacent to each other in the tumor sample. Thus, the two slices of the tissue are substantially similar. One of the pair of slides is stained only with H&E, and another separate slide of the pair of slides is stained with IHC staining of a specific molecular target. The area on the H&E stained slide corresponding to the area where IHC staining appears in the pair of slides is annotated by a human analyst as containing the specific molecular target, and the tissue class locator receives the annotated H&E slide as the training set. Substantially similar slides include, for example, in the pair of slides, if the pair of slides includes an H&E stained slide and a slide formed of molecular sequence data cut from one of the adjacent slides, or if one is stained with IHC and the other is formed of molecular sequence data, or if both are formed of similar molecular sequence data, and other combinations are included.
[0309] For example, in some embodiments, two or more samples are obtained from the subject (e.g., two or more adjacent tissue slices can be taken). In some cases, the tissue slices are obtained such that a part of the pathological slide prepared from each slice is imaged, while a part of the pathological slide is used to obtain sequencing information.
[0310] To train an optimization model according to embodiments of the present disclosure, an appropriate training dataset can be used. In some embodiments, curation of the training dataset may include collecting a series of pathology reports and related sequencing information from multiple patients. For example, a physician can perform a tumor biopsy on a patient by extracting a small amount of tumor tissue / specimen from the patient and sending this specimen to a laboratory. In the laboratory, slides can be made from the specimen using slide-making techniques such as freezing of the specimen and slice layers, setting of the specimen in paraffin and slice layers, spreading of the specimen on the slide, or other methods known to those skilled in the art. For the purposes of the following disclosure, slides and slices can be used interchangeably. A slide stores a slice of tissue from the specimen and receives a label that identifies the specimen from which the slice was extracted and the sequence number of the slice from the specimen. Conventionally, a pathology slide can be made by staining the specimen to reveal cell characteristics (such as cell nuclei, lymphocytes, stroma, epithelium, or other cells, whole or in part). The pathology slide selected for staining is conventionally the end slide of the specimen block. The slicing of the specimen proceeds using a series of initial slides that can be made for staining and diagnostic purposes. A series of subsequent consecutive slices can be used for sequencing, and the last end slide can be processed for additional staining. If the end staining slide is too far from the sequenced slides, another slide close to the sequenced slides can be stained to divide the sequenced slides by the staining slide. Although there is a slight deviation for each slice, the deviation is expected to be minimal since the tissue is sliced to a thickness close to 4um for paraffin slides and 35um for frozen slides. In the laboratory, it is generally confirmed that there is no substantial deviation in the tissue slices at a distance usually less than 40um (about 10 slides / slice).
[0311] If the specimen slices vary greatly (with low frequency) from slice to slice, the outliers may be discarded and no further processing may be required. The pathological slide 510 can be various stained slides taken from a tumor sample of a patient. Some slides and sequencing data may be obtained from the same specimen to ensure the robustness of the data, while other slides and sequencing data may be obtained from their respective unique specimens. The larger the number of tumor samples in the dataset, the higher the accuracy that can be expected from the prediction of the cell type RNA profile. In some embodiments, the stained tumor slides can be reviewed by a pathologist for the identification of cell characteristics (such as the amount of cells, their type, or differences from normal cells of the same or similar types).
[0312] In this case, the trained tissue classification model 320 receives a digital image of the H&E stained tissue and predicts tiles that may contain IHC staining or a given molecular target, and the overlay map generator 326 generates an overlay map indicating which tiles are likely to contain the IHC target or a given molecule. In one example, the resolution of the overlay is at the level of individual cells.
[0313] Overlays generated by a model trained by one or more training sets can be reviewed by a human analyst to annotate the digital slide image and add it to one of the training sets.
[0314] The pixel data patterns detected by the algorithm may represent visually detectable features. Some examples of those visually detectable features may include color, texture, cell size, shape, and spatial configuration.
[0315] For example, the color of the slide provides context information. The purple area on the slide indicates a high cell density and a high likelihood of an invasive tumor. The tumor also causes the surrounding stroma to become more fibrotic in the fibrotic reaction, usually making the normally pink stroma appear bluish-gray. The color intensity also helps to identify individual cells of a specific type (e.g., lymphocytes are uniformly very dark blue).
[0316] Texture refers to the distribution of staining within the cells. Most tumor cells have a coarse and non-uniform appearance, with bright pockets and dark nucleoli in the nucleus. A zoomed-out view with many tumor cells will have this general appearance. Each of the many non-tumor tissue classes has a characteristic function. Additionally, the pattern of tissue classes present in a region can indicate the type of tissue or cell structure present in that region.
[0317] Furthermore, cell size often indicates the tissue class. If a cell is several times larger than normal cells in other locations on the slide, it is more likely to be a tumor cell.
[0318] The shape of individual cells, particularly how round they are, can indicate what type of cells they are. Fibroblasts (stromal cells) are usually long and thin, while lymphocytes are very round. Tumor cells may have a more irregular shape.
[0319] The organization of groups of cells can also indicate the tissue class. In many cases, normal cells are organized in a structured and recognizable pattern, while tumor cells grow in a more dense and disorderly cluster. Each type and subtype of cancer can produce a tumor with a specific growth pattern, including the location of cells relative to tissue features, the spacing of tumor cells relative to each other, and the formation of geometric elements.
[0320] The technology described in this specification can be extended to other architectures. FIG. 14 shows, for example, an imaging-based biomarker prediction system 1400 that similarly uses separate pipelines for tissue classification and cell classification. System 1400 can be used for the determination of various biomarkers, including PD-L1, as described in the examples of this specification. Further, similar to other architectures described in this specification, system 1400 can be configured to predict the status of biomarkers, the status of tumors, and tumor statistics based on 3D image analysis.
[0321] System 1400 receives one or more digital images of a pathological tissue slide and creates a high-density grid-based digital overlay map, by which the majority class of tissue visible within each grid tile of the image is identified. System 1400 can also generate a digital overlay drawing that identifies each cell within the pathological tissue image at the resolution level of individual pixels.
[0322] System 1400 includes a tissue detector 1402 that detects regions of a digital image having tissue and stores data including the location of the regions detected as having tissue (e.g., the location of pixels using a reference location within the image such as the location of 0,0 pixels). The tissue detector 1402 transfers the tissue region location data 1403 to a tissue class style grid projector 1404 and a cell tile grid projector 1406. The tissue class style grid projector 1404 receives the tissue region location data 1403 and performs tissue classification on the tiles for each of some tissue class labels. A tissue class locator 1408 receives the resulting tile classifications, calculates a percentage representing the likelihood that the tissue class labels accurately describe the images within each tile, and determines where each tissue class is located in the digital image. For each tile, the sum of all percentages calculated for all tissue class labels equals 1, reflecting 100%. In one example, the tissue class locator 1408 assigns one tissue class label to each tile and determines where each tissue class is located in the digital image. The tissue class locator stores the calculated percentages and the assigned tissue class labels associated with each tile.
[0323] In one example, system 1400 includes a multi-tile algorithm that analyzes many tiles within an image both individually and in combination with the portions of the image surrounding each tile simultaneously. The multi-tile algorithm can achieve multi-scale and multi-resolution analysis that captures both the content of the individual tiles and the context of the portions of the image surrounding the tiles. Since the portions of the image surrounding two adjacent tiles overlap, analyzing many tiles and their surroundings simultaneously rather than each tile individually and its surroundings reduces computational redundancy and improves processing efficiency.
[0324] In one example, the system 1400 can store the analysis results in a three-dimensional probability data array, which includes one one-dimensional data vector for each analyzed tile. In one example, each data vector includes a list of percentages that sum to 100%, each indicating the probability that each grid tile contains one of the analyzed tissue classes. The position of each data vector within the orthogonal two-dimensional plane of the data array relative to the other vectors corresponds to the position of the tile associated with that data vector within the digital image relative to the other tiles.
[0325] The cell type tile grid projector 1406 receives the tissue region location data 1403, identifies and classifies the cells within the tile, and projects the cell type tile grid onto the region of the image containing the tissue. The cell type locator 1410 can detect each biological cell in the digital image within each grid, outline the outer edge of each cell, and classify each cell by cell type. The cell type locator 1410 stores data including the location of each cell and each pixel including the outer edge of the cell, as well as the cell type label assigned to each cell.
[0326] The overlay map generator and metric calculator 1412 can retrieve the three-dimensional probability data array stored from the tissue class locator 1408 and convert it into an overlay map that displays the tissue class labels assigned to each tile. The tissue class assigned to each tile may be displayed as a transparency color unique to each tissue class. In one example, the tissue class overlay map displays the probability of each grid tile of the tissue class selected by the user. The overlay map generator and metric calculator 1412 also searches the cell type locator 1410 for the stored cell location and type data and calculates a metric related to the number of cells within the entire image or within the tiles assigned to a particular tissue class.
[0327] FIG. 15A shows an overview of an exemplary process 1500 implemented by an imaging-based biomarker prediction system 1400, which shows a model inference pipeline for predicting biomarkers in this exemplary PD-L1. In process 1500, the fully convolutional model architecture of system 1400 is utilized and many tiles are processed in parallel. In one example, when using a GeForce GTX 1080 Ti GPU and a 6th generation Intel® Core™ i7 processor, process 1500 took 2.8 seconds to classify a single 4096×4096 pixel image. Process 1500 further included tissue detection and artifact removal algorithms so as to function in a fully automated manner in an actual setting where the slides may contain artifacts.
[0328] In a first process 1502, initial tissue segmentation is performed by a tissue detector 1402. For example, by applying a tissue masking algorithm, the outline of the tissue (red contour) is automatically drawn to generate a bounding box (not shown) around the tissue of interest. Aligned with the upper left corner of the bounding box, the tissue region is divided into non-overlapping large 4096×4096 input windows (blue dashed lines). Typically, 10 to 30 input windows are required to cover the tissue. The large window area that extends beyond the boundary region is filled with 0 (gray region).
[0329] In process 1504, a trained classification model prediction is performed. In the illustrated example, a large input window contained 128×128 = 16,384 small 32×32 tiles (the grid is much finer than shown in the figure). The large input window was padded with 0s on all sides (length 217), considering the edges of overlapping 466×466 tiles centered on each 32×32 small tile. Each large window passed through one or more trained models 1506 of a deep learning framework (including the tissue classification process of projector 1404 and the cell classification process of projector 1406). If the trained model 1506 is fully convolutional, in one example, each tile within the large input window is processed in parallel, and a 128×128×3 probability cube (where there are 3 classes) is generated. Each 1×1×3 vector of this probability cube corresponds to a 32×32 pixel region centered on the center of each 466×466 tile of the original image. The resulting probability cube is assembled into a probability map for the entire image.
[0330] Process 1508 displays an image related to the tissue masking step of process 1502. The assembled probability map generated by process 1504 passes through this tissue mask, and the background is removed. In this example, both the background region and the marker region are removed by the masking algorithm of process 1508.
[0331] Process 1510 displays an image showing a classification map that identifies one or more classified regions, such as identifying each different region of biomarker classification. In one example, the maximum probability class (argmax) is assigned to each tile through process 1508, and classification maps for 3 biomarker classes (PD-L1+, PD-L1-, other) are generated in process 1510. The classification maps show each of these biomarker classifications and their identified locations corresponding to the original pathological tissue image.
[0332] Process 1512 performs a statistical analysis on the biomarker classification from the classification map and displays the prediction scores obtained as a result of the biomarker. In this example, the number of predicted PD-L1 positive tiles is divided by the total number of predicted tumor tiles to achieve an exemplary model score.
[0333] The tissue class locator 1408 may have an architecture such as architecture 1200, for example. The architecture approximates the architecture of the FCN tile resolution classifier. Architecture 1550 can be formed from three main components: 1) a fully convolutional residual neural network (e.g., built on ResNet-18) backbone that processes a wide field of view image (FOV), 2) two branches that process small FOVs, and 3) the concatenation of small FOV features and large FOV features for multi-FOV classification. The ResNet-18 backbone includes a plurality of shortcut connections indicated by dotted lines, and the feature map is also downsampled by a factor of 2. The small FOV branch appears after the second convolutional block. The feature map of the small FOV branch is downsampled by a factor of 8 to match the dimensions of the ResNet-18 feature map. These feature maps are concatenated before passing through the softmax output to generate a PD-L1 biomarker prediction (confidence) map.
[0334] In one example, the backbone of the model includes an 18-layer version of ResNet (ResNet-18) with some modifications. The ResNet-18 backbone was converted to a fully convolutional network (FCN) by removing the global average pooling layer and eliminating zero-padding of the downsampled layers. This enabled the output of a 2D probability map instead of a 1D probability vector (see Figure 15B). In the illustrated example, the tile size (466 × 466 pixels) exceeds twice the tile size of the standard ResNet, providing the model with a large field of view (FOV) capable of learning surrounding morphological features. The tissue class locator 1408 in this example adds multiple fields of view (multi-FOV) functionality to the ResNet architecture, but it should be understood that the tissue class locator 1408 can be composed of a separate network architecture adapted to incorporate the multi-FOV approach disclosed herein.
[0335] The FCN configuration of the architecture 1505 provides many advantages, such as overcoming the problem of accuracy degradation that conventional "very deep" neural networks (including neural networks with more than 16 convolutional layers) have had (see, for example, "Deep Residual Learning for Image Recognition" by He et al. (2015) (arXiv ID: 1512.03385v1) and "Very Deep Convolution Networks for Large-Scale Image Recognition" by Simonyan et al. (2014), arXiv ID: 1409.1556v6). The architecture 1550 includes a stack of convolutional layers interleaved with "shortcut connections" that skip intermediate layers. These shortcut connections use past layers as reference points to learn the residual between layer outputs, guiding deeper layers rather than learning the identity mapping between layers. This innovation improves the convergence speed and stability during training and the performance of deeper networks compared to shallower networks.
[0336] The tissue class locator 1408 may include two additional branches with receptive fields restricted to a small FOV (32×32 pixels) centered on the second convolutional feature map (see Fig. 15B). One branch passes a copy of the small FOV to the convolutional filter, and the other branch is a standard shortcut connection using downsampling. The features generated by these additional branches are concatenated with the features from the main backbone in the softmax layer just before the output of the model is converted to probabilities. In this way, the tissue class locator 1408 combines information from multiple FOVs so that a pathologist depends on various zoom levels when diagnosing a slide. Furthermore, this architecture ensures that the central region of each tile contributes more to classification than the edges of the tile, resulting in a more accurate classification map across the whole pathological tissue image.
[0337] Fig. 15B shows an exemplary training process 1570 for generating an imaging-based biomarker prediction system 1400 and an overlay map output, where the location of the PD-L1 biomarker is predicted from the analysis of IHC and H&E pathological tissue images. In the model training process 1570, annotations by medical experts were attached to the matching regions of the IHC and H&E digital images. However, in some examples, the staining of adjacent tissue slices may be automatically annotated. Images indicating PD-L1+ and PD-L1- were annotated and used for training the system. The annotated regions of the H&E images were tiled into tiles (466×466 pixels) overlapping with a 32-pixel stride to generate training data. Next, the tissue class locator 1408 was trained using a cross-entropy loss function. The yellow squares in the model circuit diagram indicate the central regions trimmed for the small FOV. The resulting PD-L1 classification model is saved in the trained deep learning framework 1574.
[0338] FIG. 15B also shows an exemplary prediction process 1572 in which each image is divided into large non-overlapping 4096×4096 input windows (blue dashed lines). Each large window passed through the trained model. Since the deep learning framework 1574 is fully convolutional, each tile within the large input window was processed in parallel, generating a probability cube of 128×128×3 (the last dimension represents three classes). The resulting probability cube was slotted into place and assembled to generate a probability map for the entire image. The class with the highest probability was assigned to each tile, generating a PD-L1 prediction report.
[0339] FIGS. 16A-16F show an input pathologic tissue image received by an imaging-based biomarker prediction system 1400, a corresponding overlay map generated by the system 1400 to predict the location of the IHC PD-L1 biomarker, and a corresponding IHC stained tissue image used as a reference to determine the accuracy of the overlay map. The IHC stained tissue images were obtained from the test cohort but not applied to the system 1200 during model training. FIGS. 16A-16C show examples of representative PD-L1 positive biomarker classifications. FIG. 16A shows the input H&E image, FIG. 16B shows the probability map overlaid on the H&E image, and FIG. 16C shows the PD-L1 IHC staining for reference. FIGS. 16D-16F show examples of representative PD-L1 negative biomarker classifications. FIG. 16D shows the input H&E image, FIG. 16E shows the probability map overlaid on the H&E image, and FIG. 16F shows the PD-L1 IHC staining for reference. The color bar indicates the predicted probability of the tumor PD-L1+ class.
[0340] Among the advantages provided by the deep learning framework of this specification, the accuracy can be improved by breaking the shift invariance. Shift invariance, that is, uniformity, is a characteristic of linear filters such as convolution, and the response of the filter does not explicitly depend on the position. That is, when the signal is shifted, the output image is the same but the shift is applied. Although shift invariance is desirable in most image classification tasks (Le Cun (1989)), in the examples of this specification, it is generally not desirable for objects near the edge of the tile to contribute equally to the classification.
[0341] FIG. 17 shows exemplary advantages of the multi-FOV strategy as described with reference to FIGS. 14, 15A, and 15B. The upper large FOV (red box) contains both PD-L1+ tumor cells (purple, upper left) and stroma (pink). Only the stroma is included in the small FOV (green box). After passing through the convolutional layer of the tissue class locator 1408, the tumor area generates a unique pattern (colored square) that is different from the pattern generated by the stroma area (white square). After the patterns from the large and small FOV branches are concatenated, the model may predict "others". In the lower part, the field of view is shifted and the PD-L1+ tumor area fits within the small FOV. This tumor area generates the same convolutional filter pattern in both the large FOV branch and the small FOV branch (colored square). When the learned features are concatenated, the tissue class locator 1408 is more likely to predict a PD-L1+ tumor. Therefore, without the small FOV used for training the tissue class locator 1408, the system 1400 might have predicted a PD-L1+ tumor for both images. Instead, with the multi-FOV strategy of the architecture of FIG. 15, the network can preferentially classify what is in the center of the image while utilizing the rich contextual information of the surrounding area.
[0342] Still other architectures can be used with any of the classifiers herein to predict biomarker status, tumor status, and / or their metrics, especially using multi-instance learning techniques.
[0343] In the examples discussed herein, for example, a classification model architecture based on the FCN architecture described in FIG. 12A may be trained on a digital image of a pathological tissue slide that may include a matrix of annotations. Training from such digital images is performed tile by tile. For example, only tiles with annotations (i.e., labels) are provided to the deep learning framework as training tiles. Tiles of digital images without annotations may be discarded. Further, each column and row of the matrix corresponds to a separate grid having NxM pixels of the digital image. In some examples, to appropriately assign an annotation to a digital image having multiple tiles from a matrix having columns and rows, column (i) and row (j) are obtained from the matrix, and the annotation is assigned to the central region of the grid starting with pixels N(i) and M(j) and extending to the next [N-1] to [M-1] pixels. Here, the annotation of the matrix of i, j is assigned to the tile whose central region spans its range. Thus, the matrix can accurately represent the annotation for each tile within the larger digital image.
[0344] The FCN architecture can use large tiles as input while the labels are being retrieved from the central region where they are mapped to annotation mask points. The FCN architecture can learn from both the central region of the large tile and the pixels surrounding the central region, and the central region contributes more to the prediction. Further, the slide metadata can be stored in a feature vector such as a vector associated with the slide that identifies slide-level labeling. This includes patient characteristics that may improve the performance of the model. A plurality of tiles from the grid of the digital image and the corresponding annotation matrix are sequentially provided to the FCN architecture, and the FCN architecture is trained to classify the tiles of size N×M according to the annotations included in the matrix and to classify the slide itself according to the annotations included in the feature vector. The output of the FCN architecture may include a matrix in which the classification is predicted for each tile, and can be aggregated into a vector in which the classification is predicted for each slide. The matrix can be converted into a digital overlay by associating the highest classification of each tile with a color that can be overlaid at the corresponding grid position within the digital image. In some examples, the matrix can be converted into a plurality of digital overlays, each overlay corresponding to each classification, and the intensity of the associated color is assigned to the overlay based on the percentage of confidence associated with each classification. For example, a tile with a 30% probability of being a tumor, a 50% probability of being stroma, and a 20% probability of being normal can be assigned a single overlay of stroma with the highest likelihood of the tissue within the tile, or the first overlay with a first color intensity of 30% and the second overlay with a second color intensity of 50% can be assigned to identify the types of tissue that the tile may consist of.
[0345] However, even a classification model based only on architectures that may not support per-tile annotations, such as Resnet-34 or Inception-v3, can be trained using only digital images of pathological slides that contain only the vector of annotations. Here, each entry of the vector is an annotation of the patient's characteristics or metadata applied to the slide. In some examples, even an architecture that supports per-tile annotations may not have access to per-tile annotations for specific features.
[0346] To use pathological tissue images in the training of neural network images without tile annotations, or to train a neural network to identify biomarkers trained on molecular training data, in some examples, the present technology includes a deep learning training architecture configured for label-free training. For example, when training a tissue classification model, there are architectures that do not require per-tile annotations. Further, the label-free training architecture is independent of the neural network in that it enables label-free training independent of the neural network configuration (i.e., ResNet-34, FCN, Inception-v3, UNet, etc.). The architecture can analyze a set of possible training images and predict tiles to be excluded from training. Thus, in some examples, an image may be discarded from training, while in other examples, a tile may be discarded, but the remaining part of the image may be used for training. These techniques significantly reduce the amount of training data and the time required for training, and in some cases, significantly reduce the time required for updating the training of the classification model herein. Further, since this technique does not require labeling by a pathologist, the time required for training the classification model is significantly reduced, and annotation errors and annotation variations among experts are avoided.
[0347] Instead, in some examples, training can be performed with weak teacher learning that includes only image-level labeling and does not include local labeling such as tissue, cells, tumors, etc. The architecture can be composed of a label-free training front-end with an algorithm having a customized cost function, and the cost function selects the tiles to be used as input for a particular label. This process may be iterative. First, each pathological tissue image is treated as a set of tiles, and a single label for the image is applied to all tiles within the set. The tiles can be applied to an inference pipeline, such as via a network like ResNet34, Inception-v3, or FCN. Also, using pre-defined tile selection criteria such as the probability of the neural network output, output image tiles can be selected to be provided as input to the same neural network for the next round. This process can be repeated multiple times, and as sufficient sets and tiles are provided as input to the neural network, different classes of tiles are learned to be distinguished with higher accuracy as more iterations are performed.
[0348] Figure 18 shows an exemplary machine learning architecture 1800 in an exemplary configuration for performing label-free annotation training of a deep learning framework to execute the processes described in the examples of this specification. The deep learning framework 1802 can be similar to other deep learning frameworks described herein having multi-scale and single-scale classification modules, and includes a preprocessing and postprocessing controller 1804, and the execution process is described in similar examples of FIGS. 1 and 3. The deep learning framework 1802 includes a cell segmentation module 1806 and a tissue classifier module 1808, each of which is configured as a tile-based neural network classifier. The deep learning framework 1802 further includes a plurality of different biomarker classification models 1810, 1812, 1814, and 1816, each of which can be configured to have a different neural network architecture, some of which may have a multi-scale configuration and some of which may have a single-scale configuration. These different neural network architectures can be configured for training using annotated or unannotated images. Some of these architectures can be configured for training using training images with tile annotations, while other architectures of these can be configured for training using training images without tile annotations. For example, some architectures may be configured to accept only slide-level annotations of the image (i.e., annotations of the entire image, not annotations identifying characteristics of specific tiles, cell segmentation, or tissue segmentation). Examples of neural network architecture types for modules 1810-1816 include ResNet-34, FCN, Inception-v3, and UNet.
[0349] The annotated image 1818 can be provided to the deep learning framework 1802 for training various classification modules using the above-described techniques. In some examples, the entire pathological tissue image is provided to the framework 1802 for training. In some examples, the annotated image 1818 is passed directly to the deep learning framework 1802. In some examples, the annotated image 1818 can be annotated at a granularity to be reduced. Thus, in some examples, the multi-instance learning (MIL) controller 1821 can be configured to further separate the annotated image 1818 into a plurality of tile images, each corresponding to a different portion of the digital image, and the MIL controller 1821 applies those tile images to the deep learning framework 1802. However, in the architecture 1800, the unannotated images 1820 can be used for training the classification module by first providing those images 1820 to the MIL controller 1821 having a front-end tile selection controller 1822. In some examples, the MIL controller 1821 can be configured to separate the unannotated image 1820 into a plurality of tile images, each corresponding to a different portion of the digital image, and the MIL controller 1821 applies those tile images to the deep learning framework 1802. In one example, the architecture 1800 deploys weak teacher learning to train convolutional neural network architectures (such as FCN, ResNet34, Inception-v3, etc.) and classifies local tissue regions using only slide-level labels. Since location-specific annotations are not required in weak teacher learning, labeling can be performed faster, and as a result, the set of labeled slides can be increased. Thus, this architecture 1800 can be used to train a model that complements or improves FCN classification or to further train the FCN-based model itself.
[0350] In this illustrated example, the front-end tile section controller 1822 is configured in a feedback configuration that enables combining the tile section process with a classification model, allowing the tile section process to be informed by a neural network architecture such as an FCN architecture. In some examples, the tile selection process executed by the controller 1822 is a trained MIL process. For example, one of the biomarker classification models 1810 - 1816 can generate an output during training, and that output is used as an initial input to guide the tile selection controller 1822. The MIL process is typically an iterative process where the initial tile selection is difficult. For example, by informing the MIL process of the controller 1822 with guidance from the FCN architecture prediction, the MIL process of the controller 1822 can start with better examples and converge more quickly to a stable and useful FCN classifier. In yet another example, combining the results from an FCN (or other neural network) architecture and the MIL process of the controller 1822 can include combining only the results of the vector outputs in the concatenation layer such that the row-column output votes optimally. The FCN architecture and the MIL process of the controller 1822 can constitute the same prediction task (i.e., look for the same biomarker). However, in some cases, the prediction outputs of the two classification processes may be different, and in such cases, combining the results by voting for the best one can be performed. In yet another example, the output from the MIL framework and the FCN architecture can be combined to obtain a better slide-level prediction. In this case, the MIL uses a slide-level loss function as a learning criterion and uses the output from the FCN architecture to derive a guided truth for the MIL loss calculation.
[0351] The tile section controller 1822 can be implemented in a plurality of different ways.
[0352] An example of the tile selection process of controller 1822 is described with reference to a basic framework of a single class. In the example of a single class, tiles of a pathological tissue image are classified as belonging to the target class (class 1) or not belonging to it (class 0). Class 0 can be regarded as the background and can be regarded as something that does not belong to the target class. In the instance-based MIL process, it is necessary to select tiles to be used as examples for training. In the case of a single-class problem, controller 1822 can be composed of a trained model that returns the following classification. That is, if there are no tiles belonging to the target class on the slide, all tiles need to return a low inference score of zero, and during training, that slide is labeled as class 0, and if there are tiles belonging to the target class on the slide, that slide is labeled as class 1, and the tiles belonging to the target class need to return a high inference score of 1, and all other tiles need to return a low inference score of 0. To train a classifier model (for example, models 1810 - 1816), it is necessary to identify tiles representing the class at the slide level. Since it is known that there are no class 1 tiles on a class 0 slide, any tile can be used as an example for training. However, the best choice is to use the tiles with the lowest model performance. Since the inference scores of all tiles need to be 0, it is necessary to train the model using the tiles with the highest scores. These tiles with the highest scores are referred to as the "top k" tiles. Here, k is an integer (for example, 5, 10, 15) indicating the number of tiles used for training from that slide.
[0353] For class 1 slides, the controller 1822 may identify the tiles that are most likely to be of class 1. This determination can be complex because class 1 slides may contain both class 0 and class 1 tiles. However, in a classifier model trained on class 0 tiles of class 0 slides, this means that the inference scores of tiles in class 1 slides that are similar to class 0 tiles will be low. Similarly, this means that the scores of tiles that are not similar to class 0 tiles need to be high. Therefore, the tiles that are actually most likely to be of class 1 are the tiles with the highest inference scores. Therefore, as an example for training the model, it is necessary to reselect the top k tiles.
[0354] This means that, as seen in the tile selection framework 1900 of FIG. 19, for both class 0 and class 1 slides, it is necessary to use the tiles with the top k scores as training examples. The numbered rows represent the inference scores of various tiles. First, an annotated histopathological image other than the tiles is provided at 1902, and model inference is performed on each of the tiles in the image in process 1904.
[0355] The framework 1900 can be used to calculate the class prediction scores of all tiles of all slides by performing model inference 1904, and the tiles with the highest scores (the top k tiles labeled 1906) from each slide are selected for training the model. The tiles 1906 are given the same label as the slide-level label received at 1902 (e.g., tiles of a class 0 slide are given a class 0 label). After selecting tiles from all slides, the model is trained in a single epoch (or iteration) in process 1908.
[0356] At the end of a training epoch (where a tile is used once to update the model's weights), the framework 1900 is used again to compute new prediction scores for all tiles of all slides. The new prediction scores are used to identify new tiles to use for training. The process of scoring these tiles and selecting the top-k tiles for training is repeated until the model reaches a stopping criterion (e.g., convergence of performance on a held-out set of slides used for validation), as determined by process 1910.
[0357] Using a weak teacher learning configuration such as 1800 of FIG. 18 has several advantages. Strong teacher annotations and local annotations where a pathologist manually marks examples of tissue classes across an entire image are costly and inefficient. Architecture 1800 can train a model using a single label on a dataset where the slides are orders of magnitude more numerous. Further, there may be classification targets that cannot be annotated. For example, currently, genetic mutations discovered by genotype biomarkers may be correlated with tissue characteristics, but what these tissue characteristics are has heretofore been unknown. However, in the present technology, genotype biomarkers can be used as slide-level labels, and using the training framework of architecture 1800, the tissue morphologies that can be used to predict the presence of a genotype can be identified. Conventional RNA / DNA analysis to predict a genotype can take weeks, but image classification to predict a genotype using the present technology can be performed in hours and on more training sets.
[0358] To avoid overfitting situations, in some examples, the tile selection controller 1822 can be configured to perform randomized tile selection. For example, in the case of training using a small dataset (<300 slide images), a framework such as the framework 1900 of FIG. 19 may overfit to some tiles of the class 1 slides. This situation can occur because if the tiles of class 1 are used for training, their scores will be high in the next epoch. Therefore, the tiles of class 1 are likely to be selected again for training in the next epoch, the inference score will further increase, and it will be more likely to be selected again.
[0359] Figure 20 shows a framework 2000 that can be used to avoid overfitting. The framework 2000 is similar to the framework 1900 and has similar reference numbers, but uses a random tile selector 2012 to randomly select from among the high-score tiles, and then the randomly selected tile is sent to the training process 2008. The model 2004 continues to be used to determine the inference score, and the tiles are selected based on the score. However, if the inference score of a tile is high, even if it is not one of the top k tiles, it may be considered likely to be a class 1 tile. For example, the framework 2000 can set a low threshold score of 0.9 (or any value), and then any tile with a score of 0.9 or higher can be used as a training example. That is, any of the tiles 2006 can be used for training because their scores exceed the determined threshold (e.g., threshold 0.9). Next, the random high-score tile selector 2012 randomly determines which of these tiles are provided to the model training process 2008. In other examples, the tiles sent to the random high-score tile selector 2012 are the top k tiles. Additionally, in some examples, the tile selection probability applied by the selector 2012 can be completely random across all tile scores, and in other examples, the selection probability can be partially random, with tiles within a specific score or specific score range having a different random selection probability than tiles within other scores or other score ranges.
[0360] Figure 21 shows another framework 2100 for dealing with the situation of overfitting. In some examples, when training with a smaller dataset (<300 slide images), the framework of Figure 19 may overfit the slide images by predicting all tiles of a class 0 slide as class 0 and all tiles of a class 1 slide as class 1 (however, not all tiles of a class 1 slide are actually class 1). This means that there is a possibility of misclassifying the tiles of a class 1 slide. However, in framework 2100, similar to framework 2000, random tile selection is performed on high-score tiles 2101, and furthermore, random tile selection is also performed on low-score tiles 2103. In the illustrated example, the random low-score tile selector 2102 feeds a class 0 model training process 2104. The random high-score tile selector 2106 is provided to a class 1 model training process 2108.
[0361] The examples of Figures 19 - 26 are described in the context of single-class training. The label-free training of the present technology can also be used for multi-class training. In the case of a multi-class problem, the set of slide-level labels can be set so that there are no slides with a class 0 label. For example, when training a model to predict the consensus molecular subtype (CMS), in colorectal cancer (CRC), the CMS classes are used to guide targeted treatment, but only genotype biomarkers, i.e., mutations in RNA data, are used. However, the present technology predicts the genotype through imaging, thereby eliminating the need to wait weeks to perform RNA analysis and enabling the start of targeted treatment within hours to test patients. Although the example refers to CMS, the other image-based biomarker classifier modules in this specification can be trained with an architecture 1800 including PD-L1, TMB, etc.
[0362] One problem with multi-class training is that not all classes are included in all images. For example, in the case of CMS, the labels available for training images are CMS1, CMS2, CMS3, or CMS4. However, not all tiles in each image necessarily contain biomarkers indicating any of these four classes. As a result, there may be situations where there are no slides of class 0 that can be used to identify tiles of class 0, and the trained model may misclassify tissue types with no predicted values, potentially reducing the accuracy of the model. Furthermore, slides may contain features representing classes other than the slide-level labels. For example, in CMS, there may be two or more subtypes of CRC. This means that a particular image may be labeled as CMS1 because it is the main subtype of that sample, but the entire slide image may potentially contain CMS2 tissue.
[0363] To achieve multi-class training, in some examples, the tile selection controller 1822 of FIG. 18 is configured to perform a series of processes for identifying tissue feature portions correlated with different class subtypes, such as four CMS subtypes. First, the tile selection controller 1822 can be trained by identifying only examples within positive classes. Second, the tile selection controller 1822 can apply model training by identifying positive in-class tiles and negative out-of-class tiles based on low positive in-class scores. Third, the tile selection controller 1822 can apply model training by identifying positive in-class tiles and negative out-of-class tiles based on high negative scores. Next, examples of each process will be described with reference to the training of a CMS biomarker classifier as shown in FIGS. 22-26.
[0364] Figure 22 shows a framework 2200 that can be used to identify the CMS class to which each tissue feature part is most correlated, regardless of whether the tissue feature part is related to classification. Figure 22 shows an example of a first process that identifies only examples within the positive class and shows only two classes for simplicity. For each tile of the pathological tissue image, a list of possible scores is shown, and each row corresponds to the class (0, 1, and 2) shown on the left. As shown, for the slide image of class 1 (left side of Figure 22), tiles with high scores for class 1 (shaded) are used for training, and for the slide image of class 2 (right side of Figure 22), tiles with high scores for class 2 (shaded) are used for training.
[0365] In this example, all tiles are classified as 1 or 2, and the probability of being classified into class 0 becomes 0.00. Figure 23 shows an overlay map of the results showing the classification of the CMS biomarker with four classes. The tiles are color-coded based on the inference scores of the four CMS classes, where CMS1 (microsatellite instability immune) is shown in red, CMS2 (epithelial gene expression profile, WNT and MYC signaling activation) is shown in green, CMS3 (epithelial profile with obvious metabolic dysregulation) is shown in dark blue, and CMS4 (mesenchymal, significant transforming growth factor-β activation) is shown in light blue. In the illustrated example, the transparency of the tiles is adjusted to show the inference scores, with higher scores being more opaque and lower scores being more transparent.
[0366] FIG. 24 shows the framework 2200 applied in the second process. That is, model training is continued by identifying positive in-class tiles and negative out-of-class tiles by low positive in-class scores. The first process of FIG. 22 classifies all tiles as one of the non-zero classes, and the second process of FIG. 24 identifies tiles likely to be of the background class 0. In the illustrated example, the process does this by identifying tiles with scores below the slide-level class threshold. If the score of a tile in the slide image of class 1 is less than 0.1, the low score is marked as a slide image that can be used as an example of class 0. A similar process is executed for the slide image of class 2. FIG. 25 shows the overlay map obtained as a result of the classification of the CMS. If a tile is predicted to be of class 0, the tile becomes transparent. Compared with FIG. 23, in FIG. 25, this second process identifies some tissue types as class 0. This means that the image is background tissue that does not correlate with any of the CMS classes. At the same time, in the second process, various tissue types or tissue characteristics associated with each of the four classes can be identified. The tiles are colored based on the inference scores of the four CMS classes as in FIG. 23, but for tiles where no color is displayed, they are predicted to be tiles of class 0 with no class prediction value.
[0367] FIG. 26 shows a framework 2200 applied by continuing model training by identifying positive in-class tiles and negative out-of-class tiles with a high negative score. As described above, in the case of CMS biomarkers, the CMS classes of the training images are not mutually exclusive. It is possible to include a tissue type or tissue feature part of CMS2 in a slide image of CMS1. Thus, in some examples, some tiles may be mislabeled if the tile selection controller 1822 continues to label tiles as class 0 due to a low score of the slide-level class. From the second process described above, class 0 tissue that has little or no correlation with any of the CMS classes has already been identified. There are tiles with a high class 0 score. Thus, in the third process shown in FIG. 26, class 0 tiles can be identified based on having a high class 0 score.
[0368] Returning to FIG. 18, using architecture 1800, any one of biomarker classification models 1810-1816 can be trained to classify different biomarkers. The CMS is described with reference to FIGS. 22-26 as an example. Further, architecture 1800 may be agnostic to the convolutional neural network configuration, that is, each module 1810-1816 may have the same or different configurations. In addition to the FCN architecture of FIGS. 10A-10C, modules 1810-1816 can be configured with a ResNet architecture as shown in FIG. 27. The ResNet architecture provides skip connections, which help avoid the vanishing gradient problem during training and the degradation problem of large-scale architectures. Pretrained ResNet models of various sizes can be used to initialize the training of new models including, but not limited to, ResNet-18 and ResNet-34. In addition to ResNet, using architecture 1800, a neural network with convolutional layers that create feature maps input to a fully connected classification layer such as AlexNet or VGG, a network with hierarchical modules that apply multiple convolutional kernels simultaneously such as Inception v3, MobileNet, SqueezeNet, MNASNet, etc., networks designed to require fewer parameters to increase processing speed, or a customized architecture designed to more appropriately extract relevant pathological features using a neural architecture search network such as NASNet can be trained.
[0369] By using the tile-based training architecture 1800 with a tile selection controller that does not depend on the neural network configuration, a deep learning framework can be created that can identify biomarkers that are inaccurate or require a more timely process with a single architecture (such as FCN).
[0370] Furthermore, since the tile-based training architecture 1800 has a feedback configuration, in some examples, the deep learning framework 1802 can classify regions within the training images using, for example, an FCN architecture and feedback those classified images to the tile selection controller 1822 as weakly supervised training images. For example, using an FCN architecture, specific tissue regions (e.g., tumors, stroma) can first be identified and then used as inputs to a weakly supervised training pipeline such as MIL. The model trained by the weakly supervised pipeline can be used in combination with the FCN architecture to verify or improve the results of the FCN architecture or to detect new features to complement the FCN architecture.
[0371] Furthermore, there are unannotatable biomarkers such as, for example, tissue features that have not been discovered so far and that correlate with genotype, gene expression, or patient metadata. In such cases, genotype, gene expression, or patient metadata can be used to create slide-level labels that are used to train a secondary model or the FCN architecture itself using a weakly supervised framework to detect new classifications.
[0372] Furthermore, the architecture 1800 can provide detection of tissues and tissue artifacts. Regions within the slide image containing tissue can first be detected for use as inputs to the FCN architecture model. Imaging techniques such as color or texture thresholding can be used to identify tissue regions, and the deep learning convolutional neural network models (e.g., FCN architecture) of this specification can be used to further improve the generalizability and accuracy of tissue detection. Color or texture thresholding can also be used to identify false artifacts within the tissue image, and weakly supervised deep learning techniques can be used to improve generalizability and accuracy.
[0373] Furthermore, the architecture 1800 can provide marker detection. A pathological tissue image may contain notations or annotations drawn by a pathologist with markers on the slide. For example, this indicates a macrodissection area where tissue DNA / RNA analysis needs to be performed. Using a marker detection model similar to the tissue detection model, the area selected by the pathologist for analysis can be identified. This further complements the data processing of training with weak teachers and separates the areas where DNA / RNA analysis resulting in slide-level labels has been performed.
[0374] In FIG. 28, process 2800 is provided to determine an immunotherapy treatment proposed for a patient using the biomarker prediction of the imaging-based biomarker prediction system 102 of FIG. 1, particularly the deep learning framework 300 of FIG. 3. First, a pathological tissue image such as a stained H&E image is received by system 102 (2802). In process 2804, each pathological tissue image is applied to a trained deep learning framework, such as one implementing one or more FCN classification configurations described herein. In process 2806, the trained deep learning framework applies the image to a trained tissue classifier model and a trained biomarker segmentation model to determine the biomarker status of the tissue regions of the image. In some examples, a trained cell segmentation classifier model is further used by process 2806. In process 2806, the biomarker status and biomarker metrics of the image are generated. As shown in FIG. 29, the output from process 2806 is provided to process 2808 and can be implemented in a tumor treatment decision system 2900 (which can be part of a genomic sequencing system, an oncology system, a chemotherapy decision system, an immunotherapy decision system, or other treatment decision system), and the tumor treatment decision system 2900 determines the tumor type based on the received data, which includes based on biomarker metrics, genomic sequencing data, etc. The system 2900, in process 2810, analyzes the biomarker status and / or biomarker metrics and other received molecular data against the available immunotherapies 2902, and the system 2900 recommends a matching list of possible tumor type-specific immunotherapies 2904 filtered from the list of available immunotherapies 2902 in the form of a corresponding treatment report.
[0375] In various examples, the imaging-based biomarker prediction system herein can be partially or fully deployed within a dedicated slide imager such as a high-throughput digital scanner. FIG. 30 shows an exemplary system 3000, which has a dedicated ultra-high-speed pathology (slide) scanner system 3002 such as the Philips IntelliSite Pathology Solution available from Koninklijke Philips N.V. of Amsterdam, the Netherlands. In some examples, the pathology scanner system 3002 can include multiple trained biomarker classification models. Exemplary models can include, for example, the model disclosed in U.S. Patent Application No. 16 / 412,362. The scanner system 3002 is coupled to an imaging-based biomarker prediction system 3004 and implements the processes described and illustrated in the examples herein. For example, in the illustrated example, the system 3004 includes a deep learning framework 3006 based on a tile-based multi-scale and / or single-scale classification module, and has one or more trained biomarker classifiers 3008, a trained cell classifier 3010, and a trained tissue classifier 3012. The deep learning framework 3006 performs biomarker and tumor classification on pathological tissue images and stores the classification data as overlay data with the original images in the generated image database 3014. For example, the images can be stored as TIFF files. However, the database 3014 can include JSON files and other data generated by the classification process herein. In some examples, the deep learning framework can be integrated wholly or partially within the scanner 3002, as shown in optional block 3015.
[0376] The generated image may be very large, and to manage this, an image management system and viewer generator 3016 are provided. In the illustrated example, system 3016 is shown as being external to the imaging-based biomarker prediction system 3004, connected by a private or public network. However, in other examples, all or part of system 3016 can be deployed within system 3004, as shown at 3019. In some examples, system 3016 is cloud-based and stores images generated from (or instead of) database 3014. In some examples, system 3016 generates a web-accessible cloud-based viewer, which enables a pathologist to access, display, and manipulate a pathological tissue image with various classification overlays via a graphical user interface. Examples thereof are shown in FIGS. 31 - 37.
[0377] In some examples, the image management system 3016 manages the reception of scanned slide images 3018 from scanner 3002, and these slide images are generated from imager 3020.
[0378] In the illustrated example, the image management system 3016 generates an executable viewer app 3024 and deploys that app 3024 to the app deployment engine 3022 of scanner 3002. The app deployment engine 3022 can provide functions such as GUI generation to enable the user to interact with the viewer app 3024, and an app marketplace that enables the user to download the viewer app 3024 from the image management system 3016 or other network-accessible sources.
[0379] FIGS. 31 - 37 show, in one example, various digital screenshots generated by the embedded viewer 3024, which are displayed in GUI format and enable the user to manipulate the displayed image to zoom in and out and to display various classifications of tissue, cells, biomarkers, and / or tumors.
[0380] Referring to FIG. 31, a GUI generation display 3100 is shown having a panel 3102 showing the entire pathological tissue image 3104 and an enlarged portion (1.3 times magnification) of that image 3104 displayed as a window 3106. The panel 3102 further includes the magnification corresponding to the window 3106 and the tumor content report. FIG. 32 shows the display 3100, but after the user has zoomed in on the window 3106, the magnification is 3.0 times. The same is true for FIG. 33, but the magnification is 5.7 times. FIG. 34 shows a drop-down menu 3108 listing a series of classifications that a user can select to generate a classification overlay map displayed on the display 3100. FIG. 35 shows the display 3100 having an overlay map showing the tumor-classified tissue, and in this example, it is shown that the tissue is divided into tiles and the tiles having the classification are shown. In the example of FIG. 35, the illustrated classification is a tumor classification. FIG. 36 shows another exemplary classification overlay mapping, which is one of cell classification, epithelium, immune, stroma, tumor, or others. FIG. 37 shows an enlarged cell classification overlay mapping showing that the classification can be displayed at a magnification sufficient to actually distinguish different cells in the pathological tissue image.
[0381] FIG. 38 shows an exemplary computing device 3800 for implementing the imaging-based biomarker prediction system 100 of FIG. 1. As shown, system 100 can be implemented on one or more processing units 3810, which can represent a computing device 3800, particularly a central processing unit (CPU), and / or on one or more graphics processing units (GPUs) 3811, including a cluster of CPUs and / or GPUs, and / or on one or more tensor processing units (TPUs) (also labeled 3811), all of which can be cloud-based. The features and functions described for system 100 can be stored on one or more non-transitory computer-readable media 3812 of computing device 3800 and then implemented. The computer-readable media 3812 can include, for example, an operating system 3814 and a deep learning framework 3816 having elements corresponding to the elements of deep learning framework 300, which includes a preprocessing controller 302, classifier modules 304 and 306, and a postprocessing controller 308. More generally, the computer-readable media 3812 can store a trained deep learning model, executable code, etc. used to implement the techniques herein. The computer-readable media 3812 and the processing units 3810 and TPU / GPU 3811 can store image data, tissue classification data, cell segmentation data, lymphocyte segmentation data, TIL metrics, and other data in one or more databases 3813 herein. The computing device 3800 includes a network interface 3824 communicatively coupled to a network 3850 for communicating to and / or from a portable personal computer, smartphone, electronic document, tablet, and / or desktop personal computer, or other computing device. The computing device further includes an I / O interface 3826 connected to devices such as a digital display 3828, a user input device 3830, etc.In some examples, as described herein, computing device 3800 generates biomarker predictions as electronic document 3815 that can be accessed and / or shared over network 3850. In the illustrated example, system 100 is implemented on a single server 3800. However, the functionality of system 100 can be implemented across distributed devices 3800, 3802, 3804, etc. connected to each other via communication links. In other examples, the functionality of system 100 can be distributed across any number of devices, including the illustrated portable personal computer, smartphone, electronic document, tablet, and desktop personal computer devices. In other examples, the functionality of system 100 can be cloud-based, such as one or more connected cloud TPUs customized to perform a machine learning process, for example. Network 3850 can be a public network such as the Internet, a private network such as a research institution or enterprise private network, or any combination thereof. Networks include local area networks (LANs), wide area networks (WANs), cellular, satellite, or other network infrastructure, whether wireless or wired. The network can utilize communication protocols including packet-based and / or datagram-based protocols such as Internet Protocol (IP), Transmission Control Protocol (TCP), User Datagram Protocol (UDP), or other types of protocols. Additionally, the network can include a plurality of devices that facilitate network communication and / or form the hardware infrastructure of the network, such as switches, routers, gateways, access points (such as wireless access points as shown), firewalls, base stations, repeaters, backbone devices, etc.
[0382] A computer-readable medium can include executable computer-readable code stored on a computer for programming a computer to implement the techniques herein (e.g., including a processor and a GPU). Examples of such computer-readable storage media include a hard disk, a CD-ROM, a digital versatile disk (DVD), an optical storage device, a magnetic storage device, a ROM (read-only memory), a PROM (programmable read-only memory), an EPROM (erasable programmable read-only memory), an EEPROM (electrically erasable programmable read-only memory), and a flash memory. More generally, the processing unit of computing device 1300 can represent a CPU-type processing unit, a GPU-type processing unit, a TPU-type processing unit, a field programmable gate array (FPGA), another class of digital signal processor (DSP), or other hardware logic components that can be driven by a CPU.
[0383] The exemplary deep learning frameworks described herein have been described as being configured with an exemplary machine learning architecture (FCN configuration), but note that any number of suitable convolutional neural network architectures can be used. Generally speaking, the deep learning frameworks described herein can implement any suitable statistical model (e.g., a neural network or other model implemented through a machine learning process) that is applied to each of the received images. As discussed herein, the statistical model can be implemented in a variety of ways. In some examples, machine learning is used to evaluate training images and a classifier is developed that correlates pre-defined features of the images to specific categories of TIL status. In some examples, the features of the images can be identified as a training classifier using a learning algorithm such as a neural network, a support vector machine (SVM), or other machine learning process. Once the classifier within the statistical model is properly trained with a series of training images, the statistical model can be used in real time to analyze subsequent images provided as input to the statistical model for predicting the status of the biomarker. In some examples, when the statistical model is implemented using a neural network, the neural network can be configured in a variety of ways. In some examples, the neural network can be a deep neural network and / or a convolutional neural network. In some examples, the neural network can be a distributed and scalable neural network. The neural network can be customized in a variety of ways, such as by providing a specific top layer such as a logistic regression top layer. A convolutional neural network can be considered a neural network that includes a set of nodes with associated parameters. A deep convolutional neural network can be considered to have a structure with multiple layers stacked. Neural networks or other machine learning processes can include various sizes, numbers of layers, and levels of connectivity. Some layers can correspond to stacked convolutional layers (optionally followed by contrast normalization and max pooling) followed by one or more fully connected layers.In the case of a neural network trained with a large dataset, the number of layers and the size of the layers can be increased by using dropout to address the potential problem of overfitting. In some cases, the neural network can be designed to refrain from using fully connected upper layers at the top of the network. By forcing dimensionality reduction in the intermediate layers of the network, a very deep neural network model can be designed while dramatically reducing the number of learned parameters.
[0384] A system for executing the methods described herein may include a computing device, and more specifically, may be implemented on one or more processing units, such as a central processing unit (CPU), and / or one or more graphics processing units (GPUs) including a cluster of CPUs and / or GPUs. The features and functions described may be stored on one or more non-transitory computer-readable media of a computing device and then implemented therefrom. The computer-readable media may include, for example, an operating system and software modules, or “engines,” that implement the methods described herein. More generally, the computer-readable media may be capable of storing the batch normalization process instructions of an engine for implementing the techniques herein. The computing device may be a distributed computing system such as an Amazon Web Services cloud computing solution.
[0385] The computing device includes a network interface communicatively coupled to a network for communicating to and / or from a portable personal computer, smartphone, electronic document, tablet, and / or desktop personal computer, or other computing device. The computing device further includes an I / O interface connected to devices such as a digital display, user input device, etc.
[0386] The functions of the engine can be implemented in distributed computing devices interconnected with each other via a communication link. In other examples, the functions of the system can be distributed across any number of devices, including the portable personal computers, smartphones, electronic documents, tablets, and desktop personal computer devices shown. The computing devices can be communicatively coupled to a network and to another network. The network can be a public network such as the Internet, a private network such as a research institution or corporate network, or any combination thereof. The network includes local area networks (LANs), wide area networks (WANs), cellular, satellite, or other network infrastructure, whether wireless or wired. The network can utilize communication protocols including packet-based and / or datagram-based protocols such as Internet Protocol (IP), Transmission Control Protocol (TCP), User Datagram Protocol (UDP), or other types of protocols. Further, the network can include a plurality of devices that facilitate network communication and / or form the hardware infrastructure of the network, such as switches, routers, gateways, access points (such as wireless access points as shown), firewalls, base stations, repeaters, backbone devices, and the like.
[0387] A computer-readable medium may include executable computer-readable code stored on a computer for programming a computer with the techniques of this specification (e.g., including a processor and a GPU). Examples of such computer-readable storage media include hard disks, CD-ROMs, digital versatile disks (DVDs), optical storage devices, magnetic storage devices, ROM (read-only memory), PROM (programmable read-only memory), EPROM (erasable programmable read-only memory), EEPROM (electrically erasable programmable read-only memory), and flash memory. More generally, the processing unit of a computing device can represent a CPU-type processing unit, a GPU-type processing unit, a field programmable gate array (FPGA), another class of digital signal processor (DSP), or other hardware logic components that can be driven by a CPU.
[0388] Throughout this specification, multiple instances may implement components, operations, or structures described as a single instance. Although individual operations of one or more methods are illustrated and described as separate operations, one or more of the individual operations may be performed simultaneously and need not be performed in the order illustrated. Structures and functions presented as separate components within an exemplary configuration may be implemented as a combined structure or component. Similarly, structures and functions presented as a single component may be implemented as separate components or a plurality of components.
[0389] Furthermore, certain embodiments are described herein as including logic or a number of routines, subroutines, applications, or instructions. These can constitute either software (e.g., code embodied on a machine-readable medium or in a transmission signal) or hardware. In hardware, routines and the like are tangible units capable of performing specific operations and can be configured or arranged in a specific manner. In an exemplary embodiment, one or more computer systems (e.g., stand-alone, client, or server computer systems), or one or more hardware modules of a computer system (e.g., a processor or group of processors), can be configured as hardware modules that operate to perform the specific operations described herein by software (e.g., an application or a portion of an application).
[0390] In various embodiments, the hardware modules can be implemented mechanically or electronically. For example, a hardware module can include dedicated circuitry or logic (e.g., a special-purpose processor such as a microcontroller, a field-programmable gate array (FPGA), or an application-specific integrated circuit (ASIC)) that is permanently configured to perform specific operations. A hardware module can also include programmable logic or circuitry (e.g., that included within a general-purpose processor or other programmable processor) that is temporarily configured by software to perform specific operations. It will be appreciated that whether to implement a hardware module mechanically, with dedicated and permanently configured circuitry, or with temporarily configured circuitry (e.g., configured by software) can be determined considering cost and time.
[0391] Accordingly, the term "hardware module" is to be understood as encompassing a tangible entity, one that is physically constructed, permanently configured (e.g., embedded in hardware), or temporarily configured (e.g., programmed) to operate in a particular manner or to perform certain operations described herein. Considering embodiments in which a hardware module is temporarily configured (e.g., programmed), each hardware module need not be configured or instantiated at any given instance. For example, if a hardware module includes a general-purpose processor configured using software, the general-purpose processor can be configured as different hardware modules at different times. Thus, software can configure the processor, for example, to constitute a particular hardware module at one time and a different hardware module at another time.
[0392] A hardware module can provide information to other hardware modules and receive information from other hardware modules. Thus, the described hardware modules can be considered communicatively coupled. When multiple such hardware modules are present simultaneously, communication can be achieved via signal transmission connecting the hardware modules (e.g., via appropriate circuitry and buses). In embodiments where multiple hardware modules are configured or instantiated at different times, communication between such hardware modules can be achieved, for example, via storage and retrieval of information within a memory structure accessed by the multiple hardware modules. For example, a hardware module can execute an operation and store the output of that operation in a memory device to which the hardware module is communicatively coupled. Subsequently, a further hardware module can later access the memory device to retrieve and process the stored output. A hardware module can also initiate communication with an input or output device and operate on resources (e.g., collection of information).
[0393] The various operations of the exemplary methods described herein can be performed, at least in part, by one or more processors temporarily configured (e.g., by software) to perform the relevant operations or permanently configured to perform the relevant operations. Whether temporarily or permanently configured, such processors can configure processor-implemented modules that operate to perform one or more operations or functions. The modules referred to herein can, in some exemplary embodiments, include processor-implemented modules.
[0394] Similarly, the methods or routines described herein can be at least partially processor-implemented. For example, at least some of the operations of a method can be performed by one or more processors or processor-implemented hardware modules. Certain performance of the operations can be distributed not only within a single machine but also among one or more processors deployed across several machines. In some embodiments, one or more processors can be present in a single location (e.g., within a home environment, within a workplace environment, or as a server farm), while in other embodiments, the processors can be distributed across a number of locations.
[0395] Certain performance of the operations can be distributed not only within a single machine but also among one or more processors deployed across several machines. In some exemplary embodiments, one or more processors or processor-implemented modules can be present in a single location (e.g., within a home environment, within a workplace environment, or as a server farm). In other exemplary embodiments, one or more processors or processor-implemented modules can be distributed across a number of locations.
[0396] Unless otherwise indicated, the descriptions herein using words such as "processing", "computing", "calculating", "determining", "presenting", "displaying", etc. can mean the operations or processes of a machine (e.g., a computer) that manipulates or transforms data represented as a physical (e.g., electronic, magnetic, or optical) quantity within one or more memories (e.g., volatile memory, non-volatile memory, or a combination thereof), registers, or other mechanical components that receive, store, transmit, or display information.
[0397] As used herein, any reference to "one embodiment" or "an embodiment" means that a particular element, feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment. The appearances of the phrase "in one embodiment" in various places in this specification are not necessarily all referring to the same embodiment.
[0398] Some embodiments can be described using the expressions "coupled" and "connected" along with their derivatives. For example, some embodiments can be described using the term "coupled" to indicate that two or more elements are in direct physical or electrical contact. However, the term "coupled" can also mean that two or more elements are not in direct contact with each other but still cooperate or interact with each other. The embodiments are not limited to this context.
[0399] As used herein, the terms "comprises," "comprising," "includes," "including," "has," "having," or any other variation thereof are intended to cover non-exclusive inclusion. For example, a process, method, article, or apparatus that includes a list of elements is not necessarily limited to only those elements, and may include other elements not expressly listed or inherent to such process, method, article, or apparatus. Further, unless expressly stated to the contrary, "or" means an inclusive or rather than an exclusive or. For example, condition A or B is satisfied by any one of the following: A is true (or present) and B is false (or absent), A is false (or absent) and B is true (or present), and both A and B are true (or present).
[0400] In addition, the use of "a" or "an" is employed to describe the elements and components of the embodiments herein. This is done merely for convenience and to give a general sense of the description. This description should be read to include one or at least one, and the singular also includes the plural unless it is clear that the contrary is meant.
[0401] This detailed description should be construed as merely an example, and is not intended to describe all possible embodiments, as it is not possible, even if not unrealistic, to describe all possible embodiments. Many alternative embodiments can be implemented using either current technology or technology developed after the filing date of this application.
[0402] [1] A computer-implemented method for identifying biomarkers in a digital image of a hematoxylin and eosin (H&E) stained slide of a target tissue, the method comprising: receiving the digital image at an image-based biomarker prediction system having one or more processors; performing an image tiling process on the digital image by using the one or more processors to separate the digital image into a plurality of tile images, each of the plurality of tile images including a different portion of the digital image; applying the plurality of tile images to a multi-scale deep learning framework including one or more trained deep learning multi-scale classifier models, each trained to classify a different tissue classification for each tile image, and using the multi-scale deep learning framework to determine the tissue classification of each of the plurality of tile images; using the one or more processors to identify cells in the digital image using a trained cell segmentation model; Identifying the predicted presence of one or more biomarkers associated with the digital image from the tissue classifications determined for each tile image and from the identified cells within the digital image. [2] Performing the image tiling process on the digital image includes applying a tiling mask to the digital image to separate the digital image into the plurality of tile images, according to the method of item 1. [3] The method of item 2, wherein the tiling mask includes tiles of the same size. [4] The method of item 2, wherein the tiling mask includes tiles of different sizes. [5] The method of item 2, wherein the tiling mask includes tiles having a rectangular shape. [6] The method of item 2, wherein the tiling mask includes tiles characterized by the topology and / or morphology of pixels or groups of pixels. [7] Receiving the digital image includes using the one or more processors, Obtaining the digital image at a first image resolution, Downsampling the digital image to a second image resolution, Performing luminance normalization on the pixels of the digital image, Removing non-tissue objects from the digital image, according to the method of item 1. [8] The one or more trained deep learning multi-scale classifier models, Receiving, in the multi-scale deep learning framework, a plurality of H&E slide training images from a training image dataset, each H&E slide training image having a label corresponding to a biomarker to be trained. Performing tile-based tissue classification analysis on each of the H&E slide training images, Performing pixel-based cell segmentation analysis on each of the H&E slide training images, Optionally, performing tile-based biomarker classification analysis on each of the H&E slide training images, Accordingly, further including training by generating the one or more trained deep learning multi-scale classifier models, the method according to item 1. [9] Each H&E slide training image includes a plurality of tile images each having a tile-level label, the method according to item 8.
[10] For each H&E slide training image, further including attaching a tile-level label to each of the plurality of tile images of the H&E slide training image, the method according to item 8.
[11] For each H&E slide training image, performing a tile selection process for inferring the class status of each tile image within the H&E slide training image, Based on the inferred class status, before performing the tile-based tissue classification analysis on each of the H&E slide training images, discarding tile images that do not correspond to the target class, whereby the tile-based tissue classification analysis is performed only on the selected tile images of the H&E slide training images, further including the method according to item 8.
[12] The one or more trained deep learning multi-scale classifier models, Receiving a molecular training dataset of a plurality of training tissue samples, the molecular training dataset including RNA transcr...
Claims
1. A computer-implemented method for identifying biomarkers in digital images of hematoxylin and eosin (H&E) stained slides of a target tissue, the method comprising: Receiving a digital image at an image-based biomarker prediction system having one or more processors Using one or more processors to separate the digital image into a plurality of tile images using an image tiling technique, wherein each of the plurality of tile images contains a different portion of the digital image; Using one or more processors to apply the plurality of tile images to a deep learning framework including one or more trained biomarker classification models, wherein each trained biomarker classification model is trained to classify a different biomarker, and wherein the deep learning framework includes a single-scale deep learning framework or a multi-scale deep learning framework; Using one or more processors to predict a biomarker classification for each of the plurality of tile images using one or more trained biomarker classification models; wherein one or more biomarker classification models correspond to a plurality of training tissue samples and are trained using a molecular training dataset including molecular data based on sequencing of substantially similar samples associated with each training tissue sample and including a plurality of molecular data subsets clustered by biomarker; Determining a predicted presence of one or more biomarkers in the target tissue from the predicted biomarker classifications of each tile image; and Generating a report including the digital image and a digital overlay visualizing the predicted presence of one or more biomarkers.
2. The method further comprises Receiving the digital image from a pathology slide scanner system via an electrical network, the computer-implemented method of claim 1.
3. The method further comprises receiving a molecular training dataset for a plurality of training tissue samples, wherein the molecular training dataset includes RNA transcriptome counts from sequencing of substantially the same samples associated with each training tissue sample; Perform clustering processing on the molecular training dataset to identify one or more molecular data subsets corresponding to different biomarkers, For each of the one or more molecular data subsets, receive a plurality of digital images of H&E stained training slides of training tissue samples corresponding to their respective biomarkers into an image-based biomarker prediction system having one or more processors, and Apply the plurality of tile images to a deep learning framework including one or more biomarker classification models trained to classify different biomarkers based on the plurality of digital images of the H&E stained training slides and the molecular training dataset The computer-implemented method according to claim 1 or 2.
4. Including performing a plurality of instance learning processes on the plurality of digital images of the H&E stained training slides, where the plurality of instance learning processes are as follows: Provide a tile selection process for each tile image in the H&E slide training image to assign a class status to each tile image, Based on the assigned class status, discard a subset of the tile images before applying the remaining plurality of tile images to the deep learning framework The method according to claim 3.
5. The method according to claim 3, wherein each of the plurality of digital images of the H&E stained training slides of the training tissue samples has a slide-level label.
6. The method according to claim 3, wherein each of the plurality of digital images of the H&E stained training slides of the training tissue samples is unlabeled.
7. The method according to claim 1, wherein the single-scale deep learning framework is a convolutional neural network having a ResNet configuration or an inception configuration.
8. The method according to claim 1, wherein the one or more biomarkers are selected from the group consisting of consensus molecular subtype (CMS) and homologous recombination deficiency ("HRD").
9. The method according to claim 1, wherein the digital overlay includes an overlay element for identifying the tumor content of the digital image or the tumor ratio of the digital image.
Citation Information
Patent Citations
Systems and methods for detection of biological structures and / or patterns in images
JP2017516992A
Method and system for determining the ratio of different cell subsets
JP2018512071A