Assessing feature heterogeneity in digital pathology images using machine learning techniques
A machine learning model analyzes digital pathology images to identify features and calculate heterogeneity metrics, addressing the inefficiencies of current methods and improving the accuracy and speed of cancer diagnosis and treatment planning.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-07-14
- Publication Date
- 2026-03-10
AI Technical Summary
Current techniques for identifying histology and genetic mutations in digital pathology images are time-consuming, prone to human error, and incapable of assessing feature heterogeneity, which affects the diagnosis and treatment of non-small cell lung cancers like adenosquamous carcinoma.
A computer-implemented method using a machine learning model, such as a deep learning neural network, to analyze digital pathology images by subdividing them into patches, identifying features, generating labels, and calculating heterogeneity metrics to assess tissue samples for improved diagnosis and treatment recommendations.
Automated analysis reduces human error, expedites the classification process, and provides accurate assessments of feature heterogeneity, enabling better treatment selection and clinical trial eligibility for patients with non-small cell lung cancers.
Smart Images

Figure 0007827688000001 
Figure 0007827688000002 
Figure 0007827688000003
Abstract
Description
[Technical Field]
[0001] Priority This application claims the benefit under 35 U.S.C. §119(e) of U.S. Provisional Patent Application No. 63 / 052297, filed July 15, 2020, which is incorporated herein by reference.
[0002] Technical Field The present disclosure relates generally to classifying digital pathology images and assessing the heterogeneity of features detected across slide images. [Background technology]
[0003] background Adenosquamous carcinoma of the lung has a poor prognosis compared with other non-small cell lung cancers (NSCLCs). Adenocarcinoma (ADC) and squamous cell carcinoma (SCC) are common types of NSCLC. Adenosquamous carcinoma (ASC) possesses features of both ADC and SCC in the same tumor. The incidence of ASC varies among studies but is estimated to account for 0.4 to 4% of all lung cancers. Diagnosis of these cancers depends on several factors, including adequate tumor sampling, careful examination, and objective interpretation of histologic criteria.
[0004] Certain genetic mutations are associated with NSCLC or other types of cancer. Having one or more of these mutations may influence the type of treatment a physician recommends. Therefore, identifying these different genetic mutations in a patient may impact treatment and patient outcomes. Genetic mutations commonly associated with NSCLC include tumor protein 53 (TP53) mutations, Kirsten rat sarcoma viral oncogene homolog (KRAS) mutations, epidermal growth factor receptor (EGFR) mutations, and anaplastic lymphoma kinase (ALK) mutations.
[0005] Current techniques or approaches for identifying histology (e.g., ADC cancer regions, SCC cancer regions, etc.) require manual identification in digital pathology images (e.g., entire slide images) by pathologists or other trained experts. Manual identification is time-consuming, tedious, and sometimes prone to human error. Furthermore, it is often impossible to manually identify tumor mutations from digital pathology images alone. Therefore, automated techniques or methods for identifying features, including histology, mutations, or other features of interest, in digital pathology images for NSCLC, other cancers, and other conditions are desirable. Furthermore, it is desirable to assess the heterogeneity of these features in patients with specific conditions (e.g., specific cancers), which could lead to a better understanding of tumor biology and patient responsiveness to various treatments. Summary of the Invention
[0006] Overview of Certain Embodiments In certain embodiments, a computer-implemented method includes receiving a digital pathology image of a tissue sample and dividing the digital pathology image into a plurality of patches. The digital pathology image of the tissue sample may be a whole slide scan image of a tumor sample from a patient diagnosed with non-small cell lung cancer (NSCLC). In certain embodiments, the digital pathology image or the whole slide image is a hematoxylin and eosin (H&E) stained image. For each patch, the method includes identifying image features detected within the patch and using a machine learning model to generate one or more labels corresponding to the image features identified within the patch. The machine learning model may be a deep learning neural network. In one embodiment, the image features include histology, and the one or more labels applied to the patch include cancer regions of adenocarcinoma (ADC) and squamous cell carcinoma (SCC). In another embodiment, the image features indicate genetic mutations or variants, and one or more labels applied to the patches include a Kirsten rat sarcoma viral oncogene homolog (KRAS) mutation, an epidermal growth factor receptor (EGFR) mutation, an anaplastic lymphoma kinase (ALK) mutation, or a tumor protein 53 (TP53) mutation. The method includes determining a heterogeneity metric for the tissue sample based on the generated labels. If a tissue sample is represented by patches with a mixture of different labels, it is considered heterogeneous. The heterogeneity metric can be used to evaluate the degree of heterogeneity of the identified image features and corresponding labels in the tissue sample. The method further includes generating an assessment of the tissue sample based on the heterogeneity metric. A decision regarding whether the subject is eligible for a clinical trial to test a medical treatment for a specific medical condition can be made based on the assessment. One or more treatment options may also be determined for the subject based on the assessment.
[0007] In certain embodiments, the digital pathology imaging system can output various visualizations, such as patch-based image signatures that indicate the degree of heterogeneity of identified features and corresponding labels. The patch-based signatures can be used by pathologists to visualize identified image features or to evaluate machine learning models. The patch-based signatures can also assist pathologists in diagnosing or evaluating a subject or reviewing an initial assessment. The patch-based signatures can be generated based on the identified image features and can provide a visualization of the identified image features in a tissue sample, such as displaying each of the labels corresponding to the identified image features in a different color. In one embodiment, the patch-based signature can be generated using a saliency mapping technique. In certain embodiments, the patch-based signature is a heat map that includes multiple regions. Each region of the multiple regions has an associated intensity value. One or more of the multiple regions is further associated with a predicted label for the patch of the digital pathology image.
[0008] In certain embodiments, the digital pathology image processing system can train a machine learning model (e.g., a deep learning neural network) to identify image features and generate labels corresponding to the identified image features (e.g., histology, mutations, etc.) depicted in a plurality of patches from the digital pathology images. Training the machine learning model can include accessing a plurality of digital pathology images associated with a plurality of subjects (e.g., tissue samples from NSCLC patients), identifying tumor regions in each of the plurality of digital pathology images, subdividing the plurality of digital pathology images into a set of training patches, where each training patch in the set is classified by one or more features and annotated with one or more ground truth labels corresponding to the one or more features, and training the machine learning model using the set of classified patches having ground truth labels corresponding to the features depicted in the patch. The ground truth labels are provided by a clinician.
[0009] In certain embodiments, the digital pathology image processing system can further test the accuracy of the machine learning model or validate the training and update the model based on the validation. Testing and updating the machine learning model includes accessing a specific digital pathology image of a specific subject, subdividing the specific digital pathology image into a set of patches, identifying and classifying second image features detected in the patches, generating a set of predicted labels corresponding to the identified second image features of the set of patches using the trained machine learning model, comparing the generated set of predicted labels with ground truth labels associated with the set of patches, and updating the machine learning model based on the comparison. Updating the machine learning model can include further training the machine learning model.
[0010] Using machine learning models or deep learning neural networks to classify image features (e.g., histology, genetic mutations, etc.) within digital pathology (e.g., H&E stained) images and generate corresponding labels (e.g., variant type, histological subtype, etc.) is particularly advantageous in several respects. Some of the benefits include, but are not limited to, 1) relieving users (e.g., pathologists, physicians, clinical specialists, etc.) from manually evaluating thousands of full-slide images for a study and identifying features in each of these images; 2) expediting the overall image classification and evaluation process, and once the model is sufficiently trained, it may actually reduce the likelihood of errors sometimes introduced in manual human classification and evaluation; 3) helping to identify novel, previously unknown biomarkers or features; 4) studying the role of heterogeneity in patient response to treatment; and 5) utilizing images resulting from the relatively inexpensive and rapid process of H&E staining rather than relying on expensive and time-consuming DNA sequencing for certain types of analysis.
[0011] It should be noted that the embodiments disclosed herein are merely examples, and the scope of the present disclosure is not limited thereto. Examples herein may be described with respect to specific types of cancer (e.g., lung cancer, prostate cancer, etc.). These descriptions are by way of example only, not limitation, and the techniques discussed for application to specific cancers may be applied to other types of cancer and / or other conditions without requiring significant modifications or deviations from the techniques of the present disclosure. Particular embodiments may not include all, some, or any of the components, elements, features, functions, operations, or steps of the embodiments disclosed herein. Embodiments of the present invention are disclosed in the appended claims, particularly for methods, storage media, systems, and computer program products. Any feature recited in one claim category, e.g., a method, may also be claimed in another claim category, e.g., a system. Dependencies or references in the appended claims are selected for formality reasons only. However, any subject matter resulting from an intentional reference to any preceding claim (e.g., multiple dependencies) may likewise be claimed, just as any combination of a claim and its features may be disclosed and claimed regardless of the dependencies selected in the appended claims. Claimable subject matter includes not only combinations of features as recited in the appended claims, but also any other combination of features within the scope of the claims, and each feature recited in a claim may be combined with any other feature or other combination of features within the scope of the claim. Furthermore, any of the embodiments and features described or illustrated herein may be claimed in a separate claim and / or in any combination with any of the embodiments or features described or illustrated herein or with any of the features of the appended claims. [Brief explanation of the drawings]
[0012] [Figure 1A-1B] 1 illustrates an exemplary process for identifying features in digital pathology images using machine learning models and generating a subject assessment based on heterogeneity of the identified features in the digital pathology images.
[0013] [Figure 2A-2B] 1 illustrates another exemplary process for identifying features in digital pathology images using a machine learning model and generating a subject assessment based on heterogeneity of the identified features in the digital pathology images.
[0014] [Figure 3] 1 illustrates an exemplary network including a digital pathology imaging system and a digital pathology image generation system.
[0015] [Figure 4A-4B] 1 illustrates an example visualization that can be generated based on image features and corresponding labels identified using a machine learning model.
[0016] [Figure 4C] 1 illustrates the visualization of exemplary digital pathology images using different visualization techniques.
[0017] [Figure 4D] 1 shows an example of image visualization using saliency mapping techniques.
[0018] [Figure 5A-5B] 1 illustrates an exemplary embodiment of training a machine learning model for classifying patches of digital pathology images according to the detection of image features exhibited in the patches.
[0019] [Figure 6] 10 shows exemplary experimental comparison data between predicted and true labels based on two different sampling data.
[0020] [Figure 7] 1 illustrates an exemplary method for identifying features in digital pathology images using machine learning models and generating a subject assessment based on heterogeneity of the identified features in the digital pathology images.
[0021] [Figure 8] 1 illustrates an exemplary method for training and updating a machine learning model for labeling or classifying patches of digital pathology images according to detection of image features exhibited in the patches.
[0022] [Figure 9] 1 illustrates an exemplary computer system. DETAILED DESCRIPTION OF THE INVENTION
[0023] Description of exemplary embodiments The present embodiments include automated methods for detecting various features, such as histology and mutations, in digital pathology images of samples taken from a subject. The present embodiments further include automated methods for assessing the heterogeneity of these features depicted in the digital pathology images to generate an assessment of the subject's condition, such as a diagnosis and prognosis of a particular condition, such as cancer, and a recommendation for treatment of the particular condition. In certain embodiments, an exemplary method includes using a machine learning model (e.g., a deep learning neural network) to generate slide-wide predictions of image features, generating labels for subdivided patches from the entire slide corresponding to these features, calculating a heterogeneity metric based on the labels, and generating an assessment of the subject, such as a diagnosis or prognosis, based on the calculated heterogeneity metric. In some embodiments, patch-based signatures (e.g., heat maps, statistical correlations, counts, thresholds, encodings, etc.) can be created and used to evaluate the machine learning model in identifying features within the tissue sample. The present embodiments can include using a neural network to develop such patch-based signatures and using the neural network to develop evaluation criteria for the patch-based signatures. These embodiments can help standardize and accelerate accurate identification of previously difficult subtypes and combinations of presented criteria, resulting in better targeted therapy. Furthermore, automated techniques for quantifying the relative contributions of image features (e.g., corresponding to histology) in heterogeneous tumors will lead to a better understanding of tumor biology.
[0024] In certain embodiments, training a machine learning model or neural network to classify image features and generate labels can include training the model based on image data of tissue samples from multiple patients with a particular condition, such as NSCLC patients. The training samples can be scanned at a specified resolution, and the image data can include tumor regions identified by a pathologist. Each slide tumor region can be divided into smaller image patches. For example, a patch can have an area of 512 x 512 pixels, while the original image can be approximately 100,000 x 100,000 pixels. A tissue patch classifier for identifying and classifying image features shown in the tissue patch can be developed using the whole-slide labels. The classifier can be derived from a convolutional neural network (CNN) trained using transfer learning and weakly supervised learning.
[0025] Once the machine learning model (e.g., a deep learning neural network) is sufficiently trained, the model can be applied to perform patch-level predictions on unseen test images or slides. The model can output results, including a whole-slide diagnosis performed by selecting the most common predicted features (e.g., histology) from among all extracted patches for each slide. In some embodiments, patch-based image signatures can be created to visualize or represent the degree of heterogeneity of detected features (e.g., histology) within a single tissue sample in a human-interpretable form. Embodiments can further include outputting a visualization, including a display of the patch-based image signatures, network and / or image features that contributed to the patch-level classification decision.
[0026] The use of machine learning techniques to identify features, including histology or mutations, in digital pathology images (e.g., H&E-stained images) is now described with reference to FIGS. 1A-1B and 2A-2B. In particular, FIGS. 1A-1B illustrate and describe an embodiment for using a deep learning neural network to identify or extract histology from a patch of an exemplary digital pathology image and assess the heterogeneity of the extracted histology to diagnose a subject's or patient's lung cancer condition. FIGS. 2A-2B illustrate and describe an embodiment for using a deep learning neural network to identify mutations or genetic variants from a patch of an exemplary digital pathology image and assess the heterogeneity of the identified mutations for subject assessment. Note that while some of the descriptions in these FIGS. 1A-1B and 2A-2B may overlap (e.g., processes for image classification and heterogeneity calculation), the features, labels, heterogeneity metrics, and subject assessments determined or generated in each of FIGS. 1A-1B and 2A-2B are different. Each of these figures is described in detail below.
[0027] FIGS. 1A-1B illustrate exemplary processes 100 and 150, respectively, for identifying features in a digital pathology image using a machine learning model and generating a subject assessment based on the heterogeneity of the features in the digital pathology image. Specifically, FIGS. 1A-1B illustrate using a deep learning neural network to identify or extract features, e.g., histology, from patches of an exemplary digital pathology image and assess the heterogeneity of the extracted histology to diagnose a subject's or patient's lung cancer condition. Identifying mixed histological subtypes can enable better selection of patients / subjects responsive to drug treatment and elucidate biological mechanisms leading to tumor heterogeneity. FIG. 3 illustrates a network 300 of interactive computer systems that can be used to extract histology and assess the heterogeneity of the extracted histology using a deep learning neural network, according to some embodiments described herein. As shown in FIG. 1A, at 110, a digital pathology image processing system 310 receives a digital pathology image 105 of a tissue sample. In certain embodiments, the digital pathology image of the tissue sample can be a whole-slide image of a sample from a subject or patient diagnosed with non-small cell lung cancer. In a first embodiment, the digital pathology images or whole slide images discussed herein can be H&E-stained images or can be obtained from H&E preparations of tumor samples from patients diagnosed with non-small cell lung cancer. The advantage of using H&E-stained images is that they are relatively quick and inexpensive to obtain, especially compared to other techniques such as DNA sequencing, which can be an expensive and time-consuming process for identifying specific tissue images. It should be understood that images generated by other staining and imaging techniques can also be used without departing from the teachings of the present disclosure. The digital pathology image processing system 310 can receive digital pathology images or whole slide images 105 from the digital pathology image generation system 320 or one or more components thereof. As another example, the digital pathology image processing system 310 can receive digital pathology images 105 from one or more user devices 330.The user device 330 may be a computer used by a pathologist or clinician connected via one or more networks to the digital pathology imaging system 310. A user of the user device 330 may use the user device 330 to upload digital pathology images 105 or to instruct one or more other devices to provide the digital pathology images 105 to the digital pathology imaging system 310.
[0028] In certain embodiments, although not shown in FIG. 1A , the digital pathology image processing system 310 can perform tumor lesion segmentation on the image 110. For example, tumor regions (e.g., diseased regions) can be identified in the image 110. In one embodiment, the digital pathology image processing system 310 can identify tumor regions with the assistance of a user. For example, the digital pathology image processing system 310 automatically identifies tumor regions using a separate tumor segmentation algorithm or one or more machine learning techniques. For example, a machine learning model can be trained based on pre-labeled or pre-annotated tumor regions in a set of digital pathology images by a human expert. In some embodiments, tumor regions can be manually selected by a pathologist, physician, clinical specialist, lung cancer (or another cancer of interest related to the tissue sample) diagnostic expert, etc. Performing tumor lesion segmentation is advantageous for eliminating irrelevant portions of the digital pathology image 110 (e.g., blank regions) that do not contribute to feature evaluation, ultimately reducing the prevalence of useful signals.
[0029] At 120, the digital pathology image processing system 310 subdivides the digital pathology image 105 (with the identified tumor region) into a plurality of patches 115a, 115b, 115c, ... 115n (also individually or collectively referred to herein as 115), using, for example, a patch generation module 311. Subdividing the image 105 may, in some cases, include tiling the image in a grid structure format into small image tiles or patches, as shown in FIG. 1A. Although shown in certain embodiments as occurring after tumor lesion segmentation, the tumor lesion segmentation process may occur after subdividing the image into patches.
[0030] At 130, the digital pathology imaging system 310 identifies one or more image features within each patch, for example, using a patch classification module 312, and generates, using a deep learning neural network 125, a plurality of labels 135a, 135b, 135c, ... 135n (also referred to herein individually or collectively as 135) for the plurality of patches 115 corresponding to the identified image features. In some embodiments, as described elsewhere herein, by identifying the image features, the digital pathology imaging system 310 can identify or classify tissue structures underlying the tissue sample. Each label 135 can indicate, identify, or represent a particular image feature, such as a type of non-small cell lung cancer (NSCLC). For example, for patch 115a, patch classification module 312 generates corresponding label 135a indicating that one or more image features shown in patch 115a are associated with adenocarcinoma (ADC), for patch 115b, patch classification module 312 generates corresponding label 135b indicating that one or more image features shown in patch 115b are associated with squamous cell carcinoma (SCC), for patch 115c, patch classification module 312 generates corresponding label 135c indicating that one or more image features shown in patch 115b are associated with SCC, and for patch 115n, patch classification module 312 generates corresponding label 135n indicating that one or more image features shown in patch 115n are associated with ADC. Note that although only two types of labels, ADC and SCC, are shown in FIG. 1A , other types of labels for other features in tissue structures can be identified using deep learning neural network 125.
[0031] In particular embodiments, the deep learning neural network 125 described herein is a convolutional neural network that can be trained based on Inception V3 and Resnet18 architectures using transfer learning and weakly supervised learning techniques. It should be understood that other learning techniques for training the deep learning neural network 125 are possible and within the scope of the present disclosure. Training the deep learning neural network 125 to classify image patches based on tissue features identified within the image patches is described in detail below with reference to at least FIGS. 5A-5B and 8.
[0032] FIG. 1B illustrates a process 150 for assessing the heterogeneity of identified histology in a digital pathology image (e.g., in a patch of the digital pathology image) for subject evaluation based on the labels 135 generated by process 100 of FIG. 1A. Note that, as with FIG. 1A, reference is made to FIG. 3 for the various computing components or modules that perform respective operations in process 150. As illustrated, the digital pathology image processing system 310 can optionally provide the labeled patches (e.g., patches and their corresponding histology) to any visualization tool or application 160, for example, using an image visualization module 313. For example, the image visualization module 313 can provide a patch labeled ADC 155a, a patch labeled SCC 155b, a patch labeled SCC 155c, and a patch labeled ADC 155n along with the digital pathology image 105 to the visualization tool 160. In some embodiments, the image visualization module 313 and the visualization tool 160 are combined or work together as a single component. In other embodiments, the image visualization module 313 and the visualization tool 160 are separate components.
[0033] In some embodiments, any visualization tool 160 using the labeled patches 155 and the entire digital pathology image or entire slide image 105 can generate any patch-based signature (e.g., heat map, region overlay) 170 of the digital pathology image. Note that the visualization tool 160 and patch-based signature 170 are depicted with dotted lines to indicate that they are optional parts or components of the process 150 and may or may not be used in assessing heterogeneity as described herein. In certain embodiments, the patch-based signature 170 can represent a visualization of histology within a tissue sample. The visualization may include displaying histology in different colors, as shown in FIGS. 4A-4B. By way of example and not limitation, visualizing different histologies in a tissue sample via patch-based signatures can include displaying ADC cancer regions in blue, SCC cancer regions in green, etc. In certain embodiments, the digital pathology image processing system 310, using, for example, the image visualization module 313, can generate the patch-based signatures described herein using different visualization techniques. For example, the image visualization module 313 may use one or more of a gradient-weighted class activation mapping (Grad-CAM) technique, a score-weighted class activation mapping (score-CAM) technique, an occlusion mapping technique, and a saliency mapping technique to generate the visualization. Different visualization techniques are shown and described below with reference to FIG. 4C. In a particular embodiment, the image visualization module 313 generates the patch-based signature 170 using a saliency mapping technique.
[0034] At 175, the digital pathology image processing system 310 calculates a heterogeneity metric for the histologies identified or extracted in the digital pathology image (e.g., via the patches) using, for example, the labeled patches 155 (e.g., patches and their corresponding labels) using the heterogeneity metric calculation module 314. The heterogeneity metric can assess the heterogeneity of histologies in cancer. The heterogeneity metric can include a quantifiable measure of the level or degree of heterogeneity of the histologies. In certain embodiments, the heterogeneity metric can quantify the relative proportion of each histology relative to other histologies in a given tissue sample. By way of example and not limitation, for the ADC and SCC histologies identified in FIG. 1A, the heterogeneity metric can indicate the proportion of each ADC and SCC cancer region in the tissue sample (e.g., a patient's non-small cell lung cancer image). In other words, the heterogeneity metric can indicate how many total ADC cancer regions there are relative to all SCC cancer regions in the tissue sample. For example, a heterogeneity metric may indicate that there are a total of 392 ADC and 150 SCC cancer regions in a given tissue sample. As another example, a heterogeneity metric may indicate the association between ADC and SCC cancer regions to identify how the regions are distributed in the sample.
[0035] In alternative embodiments, the heterogeneity metric calculation module 314 can calculate a heterogeneity metric based on the patch-based signature described herein. For example, the heterogeneity metric module 314 can receive the patch-based signature from the image visualization module 313 and calculate a heterogeneity metric using information indicated in the patch-based signature. As an example, the patch-based signature can indicate the distribution or proportion of each label (e.g., ADC, SCC) within a tissue sample (e.g., as shown in FIG. 4A ), and the heterogeneity metric calculation module 314 can calculate a heterogeneity metric using this distribution information in the patch-based signature. In some embodiments, the patch-based signature can correspond to a patch-level assessment of the heterogeneity of a feature (e.g., a histology image). The patch-based signature may be specific to the digital pathology image or may be understood to be related to a pattern or classification of heterogeneity of features in the digital pathology image.
[0036] The digital pathology imaging system 310 generates an output based on the heterogeneity metric, for example, using the output generation module 316. In certain embodiments, the output can include a subject assessment 180 based on the calculated heterogeneity metric. The subject assessment can include, for example, a subject diagnosis, a subject prognosis, or a treatment recommendation applicable to the operator's particular use case. For example, based on a heterogeneity metric indicating the degree to which image features (e.g., histology) and / or labels (e.g., ADC cancer regions, SCC cancer regions) are heterogeneous in a given tissue sample, the output generation module 316 can generate an appropriate assessment for the given tissue sample. As an example, the assessment can include the severity of a patient's lung cancer based on the amount of ADC and SCC cancer regions present in the patient's tissue sample. As another example, the assessment can include the best treatment option for a patient's lung cancer based on the presence or heterogeneity of ADC and SCC cancer regions present in the patient's tissue sample. In some embodiments, the output generation module 316 can provide the subject assessment 180 for display to a user, such as a pathologist, physician, clinical specialist, lung cancer diagnostic specialist, or operator of the digital pathology imaging system 310. The subject assessment 180 can also be provided to one or more user devices 330. In some embodiments, the subject assessment 180 can be used to predict a subject's responsiveness to various treatments, predict the appropriateness of one or more treatment options for the subject, identify treatments predicted to be effective for the subject, and / or assign the subject to an appropriate arm within a clinical trial. In some embodiments, the output generation module 316 can output an indication of whether the subject is eligible for a clinical trial in testing a medical treatment for a particular medical condition, based on the assessment 180.
[0037] Output from the digital pathology imaging system 310 can be provided in several forms, including a simple listing of the assessments made by the digital pathology imaging system. More advanced outputs can also be provided. As an example, the digital pathology imaging system 310 can generate different visualizations of the identified histologies discussed herein. For example, the digital pathology imaging system 310 can generate an overall map showing the various histologies, as shown in FIG. 4A. As another example, the digital pathology imaging system 310 can generate separate maps for each histology, as shown in FIG. 4B.
[0038] 2A-2B illustrate another exemplary process 200 and 250 for identifying features in digital pathology images using a machine learning model and generating a subject assessment based on the heterogeneity of the features in the digital pathology images, respectively. Specifically, FIGS. 2A-2B illustrate using a deep learning neural network to identify features, e.g., genetic mutations or variants, from patches of exemplary digital pathology images and assess the heterogeneity of the identified mutations to diagnose a subject's or patient's condition. FIG. 3 illustrates a network 300 of interactive computer systems that can be used to identify genetic mutations or variants from patches of exemplary digital pathology images using a deep learning neural network and assess the heterogeneity of the identified mutations, according to some embodiments described herein. As previously mentioned, pathologists cannot predict mutation status directly from H&E images and instead rely on DNA sequencing, which is expensive and time-consuming. Thus, using the machine learning techniques discussed herein to predict mutations or mutation status is significantly faster and more efficient than using DNA sequencing techniques.
[0039] As shown in FIG. 2A , at 210, the digital pathology imaging system 310 receives a digital pathology image 205 of a tissue sample. In certain embodiments, the digital pathology image of the tissue sample can be a whole slide image of a sample from a subject or patient diagnosed with non-small cell lung cancer. In a primary embodiment, the digital pathology image or whole slide image discussed herein can be an H&E stained image or obtained from an H&E preparation of a tumor sample from a patient diagnosed with non-small cell lung cancer. The digital pathology imaging system 310 can receive the entire digital pathology image or slide image 205 from a digital pathology image generation system 320 or one or more components thereof. As another example, the digital pathology imaging system 310 can receive the digital pathology image 205 from one or more user devices 330. The user devices 330 can be computers used by pathologists or clinicians connected to the digital pathology imaging system 310 via one or more networks. A user of the user device 330 can use the user device 330 to upload a digital pathology image 205 or instruct one or more other devices to provide the digital pathology image 205 to the digital pathology image processing system 310.
[0040] In certain embodiments, although not shown in FIG. 2A , the digital pathology imaging system 310 can perform tumor lesion segmentation on the image 210. For example, a tumor region or regions can be identified in the image 210. In one embodiment, the digital pathology imaging system 310 automatically identifies tumor regions using a separate tumor lesion segmentation algorithm or one or more machine learning techniques. In one embodiment, the digital pathology imaging system 310 can identify tumor regions with the assistance of a user. For example, tumor regions may be manually selected by a pathologist, physician, clinical specialist, lung cancer diagnostic expert, etc. At 220, the digital pathology imaging system 310 subdivides the digital pathology image 205 (with the identified tumor regions) into multiple patches or tiles 215 a, 215 b, 215 c, ... 215 n (also individually or collectively referred to herein as 215), for example, using a patch generation module 311.
[0041] At 230, the digital pathology image processing system 310 identifies one or more image features in each patch, for example using a patch classification module 312, and generates multiple labels 235a, 235b, 235c, ... 235n (also referred to herein individually or collectively as 235) for the multiple patches 215 corresponding to the identified image features using a deep learning neural network 225. Each label 235 can indicate, identify, or predict a particular mutation or genetic variant. For example, for patch 215a, the patch classification module 312 generates a corresponding label 235a indicating that one or more image features shown in patch 215a are associated with a KRAS mutation, for patch 215b, the patch classification module 312 generates a corresponding label 235b indicating that one or more image features shown in patch 215b are associated with an epidermal growth factor receptor (EGFR) mutation, for patch 115c, the patch classification module 312 generates a corresponding label 235c indicating that one or more image features shown in patch 215b are associated with a KRAS mutation, and for patch 215n, the patch classification module 312 generates a corresponding label 235n indicating that one or more image features shown in patch 215n are associated with an EGFR mutation. Note that although only two types of labels or mutations, KRAS and EGFR, are shown in FIG. 2A , other types of mutations or genetic mutations can be similarly identified for patches using the deep learning neural network 225.
[0042] FIG. 2B illustrates a process 250 for assessing heterogeneity of mutations identified in a digital pathology image (e.g., in patches of a digital pathology image) for subject evaluation based on the labels 235 generated by process 200 of FIG. 2A, as described above. Note that, as with FIG. 2A, reference is made to FIG. 3 for the various computing components or modules that perform respective operations in process 250. As illustrated, the digital pathology image processing system 310 can provide the labeled patches (e.g., patches and annotations indicating the corresponding mutations indicated) to any visualization tool or application 260, for example, using an image visualization module 313. For example, the image visualization module 313 can provide the patch labeled KRAS255a, the patch labeled EGFR255b, the patch labeled KRAS255c, the patch labeled EGFR255n, and the remaining patches along with the digital pathology image 205 to the visualization tool 260. In some embodiments, the image visualization module 313 and the visualization tools 260 are combined or work together as a single component. In other embodiments, the image visualization module 313 and the visualization tools 260 are separate components.
[0043] At 265, an optional visualization tool 260 using the labeled patches 255 and the entire digital pathology image or entire slide image 205 can generate a patch-based signature 270 of the digital pathology image. Note that the visualization tool 260 and the patch-based signature 270 are depicted with dotted lines to indicate that they are optional parts or components of the process 250 and may or may not be used in assessing heterogeneity as described herein. In certain embodiments, the patch-based signature or heat map 270 can show a visualization of mutations in a tissue sample. The visualization can include displaying predicted mutations or genetic variants with different color coding. In certain embodiments, the digital pathology image processing system 310, for example, using the image visualization module 313, can use different visualization techniques to generate the patch-based signatures described herein. For example, the image visualization module 313 can use one or more of the following techniques to generate the visualization: grad-cam, score-cam, occlusion mapping, or saliency mapping. Different visualization techniques are shown and described below with reference to FIG. 4C . In certain embodiments, the image visualization module 313 generates the patch-based signature 270 using a saliency mapping technique.
[0044] At 275, the digital pathology image processing system 310 calculates heterogeneity metrics for the mutations identified in the digital pathology image (e.g., via the patches) using the labeled patches 255 (e.g., patches and their corresponding labels), e.g., using the heterogeneity metric calculation module 314. The heterogeneity metric may include a quantifiable measure of the level or degree of heterogeneity of the mutations. In certain embodiments, the heterogeneity metric may quantify the relative proportion of each mutation relative to other mutations in a given tissue sample. By way of example, and not limitation, for the KRAS and EGFR mutations identified in FIG. 2A , the heterogeneity metric may indicate the respective percentage distributions of KRAS and EGFR mutations in a tissue sample (e.g., a patient's non-small cell lung cancer image). The heterogeneity metric may indicate the presence of regions with different mutations within the same tumor. This is important because patients who receive effective targeted therapy for tumors with a specific mutation may not respond or may later relapse if there are some areas of the tumor that do not have that mutation.
[0045] The digital pathology imaging system 310 generates an output based on the heterogeneity metric, for example, using the output generation module 316. In certain embodiments, the output can include a subject rating 180 based on the calculated heterogeneity metric. The subject rating can include, for example, a subject diagnosis, a subject prognosis, or a treatment recommendation applicable to the operator's particular use case. For example, based on the heterogeneity metric indicating the degree to which various features (e.g., mutations) are heterogeneous in a given tissue sample, the output generation module 316 can generate an appropriate rating for the given tissue sample. As an example, the rating can include appropriate treatment options for a patient's lung cancer based on the presence or heterogeneity of KRAS and EGFR gene mutations present in the patient's tissue sample. In some embodiments, the output generation module 316 can provide a subject rating 280 for display to a user, such as a pathologist, physician, clinical specialist, lung cancer diagnostic expert, or operator of the digital pathology imaging system 310. The subject rating 280 can also be provided to one or more user devices 330. In some embodiments, the subject assessment 280 can be used to predict a subject's responsiveness to various treatments, identify treatments predicted to be effective for the subject, and / or assign the subject to an appropriate arm within a clinical trial. In some embodiments, the output generation module 316 can output an indication of whether the subject is eligible for a clinical trial in testing a medical treatment for a particular medical condition, based on the assessment 280.
[0046] FIG. 3 illustrates a network 300 of interactive computer systems that can be used as described herein to identify features within patches of a digital pathology image, generate labels for the identified features using deep learning techniques, and generate an assessment based on the heterogeneity of the identified features within the digital pathology image, according to some embodiments of the present disclosure.
[0047] The digital pathology image generation system 320 can generate one or more digital pathology images, including, but not limited to, entire slide images, corresponding to a particular sample. For example, an image generated by the digital pathology image generation system 320 can include a stained section of a biopsy sample or an unstained section of a biopsy sample submitted for pre-processing. As another example, an image generated by the digital pathology image generation system 320 can include a slide image of a liquid sample (e.g., a blood film). As another example, an image generated by the digital pathology image generation system 320 can include fluorescence microscopy, such as a slide image showing fluorescence in situ hybridization (FISH) after a fluorescent probe binds to a target DNA or RNA sequence.
[0048] Several types of samples can be processed by the sample preparation system 321 to fix and / or embed the sample. The sample preparation system 321 can facilitate infiltrating the sample with a fixative (e.g., a liquid fixative such as a formaldehyde solution) and / or an embedding substance (e.g., histological wax). For example, the sample fixation subsystem can fix the sample by exposing the sample to a fixative for at least a threshold time (e.g., at least 3 hours, at least 6 hours, or at least 12 hours). The dehydration subsystem can dehydrate the sample (e.g., by exposing the fixed sample and / or a portion of the fixed sample to one or more ethanol solutions) and potentially clear the dehydrated sample using a clearing intermediate (e.g., including ethanol and histological wax). The sample embedding subsystem can infiltrate the sample with heated (e.g., liquid) histological wax (e.g., one or more times for a corresponding predetermined period of time). The histological wax can include paraffin wax and potentially one or more resins (e.g., styrene or polyethylene). The sample and wax are then allowed to cool and the wax-infiltrated sample can be blocked out.
[0049] The sample slicer 322 can receive a fixed and embedded sample and create a set of sections. The sample slicer 322 can expose the fixed and embedded sample to a cold or low temperature. The sample slicer 322 can then cut the cooled sample (or a trimmed version thereof) to create a set of sections. Each section can have a thickness of (e.g.) less than 100 μm, less than 50 μm, less than 10 μm, or less than 5 μm. Each section can have a thickness of (e.g.) more than 0.1 μm, more than 1 μm, more than 2 μm, or more than 4 μm. Cutting of the cooled sample can be performed in a warm water bath (e.g., at a temperature of at least 30° C., at least 35° C., or at least 40° C.).
[0050] The automated staining system 323 can facilitate staining of one or more of the sample sections by exposing each section to one or more stains. Each section can be exposed to a predetermined amount of stain for a predetermined period of time. In some cases, a single section is exposed to multiple stains simultaneously or sequentially.
[0051] Each of the one or more stained sections can be presented to an image scanner 324, which can capture a digital image of the section. The image scanner 324 can include a microscope camera. The image scanner 324 can capture digital images at multiple magnifications (e.g., using a 10x objective, a 20x objective, a 40x objective, etc.). Image manipulation can be used to capture selected portions of the sample at a desired magnification range. The image scanner 324 can further capture annotations and / or morphemes identified by a human operator. In some cases, the section is returned to the automated staining system 323 after one or more images are captured so that the section can be washed, exposed to one or more other stains, and imaged again. When multiple stains are used, the stains can be selected to have different color profiles so that a first region of the image corresponding to a first portion of the section that has absorbed a large amount of the first stain can be distinguished from a second region of the image (or a different image) corresponding to a second portion of the section that has absorbed a large amount of the second stain.
[0052] It will be appreciated that one or more components of digital pathology imaging system 320 may, in some cases, operate in conjunction with a human operator. For example, a human operator may move samples through various subsystems (e.g., sample preparation system 321 or digital pathology imaging system 320) and / or initiate or terminate the operation of one or more subsystems, systems, or components of digital pathology imaging system 320. As another example, some or all of one or more components of the digital pathology imaging system (e.g., one or more subsystems of sample preparation system 321) may be partially or entirely replaced by the actions of a human operator.
[0053] Furthermore, while the various described and illustrated features and components of the digital pathology imaging system 320 relate to processing solid and / or biopsy samples, it will be understood that other embodiments may relate to liquid samples (e.g., blood samples). For example, the digital pathology imaging system 320 may receive a liquid sample (e.g., blood or urine) slide including a base slide, a smeared liquid sample, and a cover. The image scanner 324 may then capture an image of the sample slide. Further embodiments of the digital pathology imaging system 320 may relate to capturing images of the sample using advanced imaging techniques, such as FISH, as described herein. For example, once fluorescent probes are introduced into the sample and allowed to bind to target sequences, appropriate imaging can be used to capture an image of the sample for further analysis.
[0054] A given sample can be associated with one or more users (e.g., one or more physicians, laboratory technicians, and / or healthcare providers) during processing and imaging. Associated users can include, by way of example and not limitation, the person who ordered the test or biopsy that produced the sample being imaged, the person authorized to receive the test or biopsy results, or the person who performed the analysis of the test or biopsy sample, among others. For example, a user can correspond to a physician, pathologist, clinician, or patient. A user can use one or more user devices 330 to submit one or more requests (e.g., to identify a subject) for a sample to be processed by the digital pathology image generation system 320 and the resulting images to be processed by the digital pathology image processing system 310.
[0055] The digital pathology image generation system 320 can transmit images generated by the image scanner 324 back to the user device 330. The user device 330 then communicates with the digital pathology image processing system 310 to initiate automated processing of the images. In certain embodiments, the images so generated after processing by one or more of the sample preparation system 321, sample slicer 322, automated staining system 323, or image scanner 324 can be H&E-stained images or images generated by a similar staining procedure. In some cases, the digital pathology image generation system 320 provides the images generated by the image scanner 324 (e.g., H&E-stained images) directly to the digital pathology image processing system 310, for example, according to instructions from a user of the user device 330. Although not shown, other intermediate devices (e.g., a data store on a server connected to the digital pathology image generation system 320 or the digital pathology image processing system 310) can also be used. Furthermore, for simplicity, only one digital pathology image processing system 310, image generation system 320, and user device 330 are shown in the network 300. This disclosure contemplates the use of one or more of each type of system and its components without necessarily departing from the teachings of the disclosure.
[0056] The network 300 and associated systems shown in FIG. 3 can be used in a variety of settings where scanning and evaluating digital pathology images, such as entire slide images, is an essential component of the work. As an example, the network 300 can be associated with a clinical environment where a user is evaluating a sample for possible diagnostic purposes. The user can review the image using a user device 330 before providing it to the digital pathology imaging system 310. The user can provide additional information to the digital pathology imaging system 310 that can be used to guide or direct the analysis of the image by the digital pathology imaging system 310. For example, the user can provide a predicted diagnosis or preliminary evaluation of features within the scan. The user can also provide additional context, such as the type of tissue being reviewed. As another example, the network 300 can be associated with a laboratory environment where tissue is being examined, e.g., to determine the efficacy or potential side effects of a drug. In this regard, it may be common to submit multiple types of tissue for review to determine the drug's systemic effect. This can present particular challenges to a human scan examiner who may need to determine various aspects of the image, which can be highly dependent on the type of tissue being imaged. These contexts can optionally be provided to the digital pathology imaging system 310.
[0057] The digital pathology image processing system 310 can process digital pathology images, including whole slide images or H&E-stained images, to classify features in the digital pathology images and generate labels / annotations for the classified features in the digital pathology images, as described above with reference to Figures 1A-1B and 2A-2B. The patch generation module 311 can define a set of patches for each digital pathology image. To define the set of patches, the patch generation module 311 can subdivide the digital pathology image into a set of patches. As implemented herein, patches can be non-overlapping (e.g., each patch includes pixels of the image that are not included in any other patch) or overlapping (e.g., each patch includes a portion of pixels of the image that are included in at least one other patch). Features such as whether patches overlap, in addition to the size of each patch and the window stride (e.g., the image distance or pixels between a patch and a subsequent patch), can increase or decrease the dataset for analysis; more patches (e.g., through overlapping or smaller patches) can increase the potential resolution of the final output and visualization, resulting in a larger and more diverse dataset for training purposes. In some cases, the patch generation module 311 defines a set of patches for an image, each patch being a predetermined size and / or with a predetermined offset between patches. Furthermore, the patch generation module 311 may create multiple sets of patches per image with varying sizes, overlaps, step sizes, etc. In some embodiments, the digital pathology image itself may contain overlapping patches that may result from the imaging technique. Even segmentation without patch overlaps may be a preferred solution for balancing patch processing requirements.The patch size or patch offset can be determined, for example, by calculating one or more performance metrics (e.g., precision, recall, accuracy, and / or error) for each size / offset and selecting the patch size and / or offset associated with one or more performance metrics above a predetermined threshold and / or associated with one or more optimal (e.g., high precision, highest recall, highest accuracy, and / or lowest error) performance metrics. The patch generation module 311 can further define the patch size according to the type of condition being detected. For example, the patch generation module 311 can be configured with knowledge of the type of histology or mutation that the digital pathology image processing system 310 will search for and can customize the patch size according to the histology or mutation to optimize detection. In some cases, the patch generation module 311 defines a set of patches, and the number of patches in each set of images, the size of the patches in the set, the resolution of the set of patches, or other related properties are defined and held constant for each of the one or more images.
[0058] In some embodiments, the patch generation module 311 can further define a set of patches for each digital pathology image along one or more color channels or color combinations. By way of example, digital pathology images received by the digital pathology image processing system 310 can include large-format, multi-color channel images with pixel color values for each pixel of the image specified for one of several color channels. Examples of usable color specifications or color spaces include RGB, CMYK, HSL, HSV, or HSB color specifications. The set of patches can be defined based on subdividing the color channels and / or generating a luminance map or grayscale equivalent for each patch. For example, for each portion of the image, the patch generation module 311 can provide a red tile, a blue tile, a green tile, and / or a luminance tile, or equivalents for the color specifications used. As described herein, subdividing the digital pathology image based on the image portions and / or color values of the portions can generate patch and image labels and improve the accuracy and recognition rate of networks used to generate image classifications. Additionally, the digital pathology image processing system 310 can convert between color specifications and / or prepare copies of patches using multiple color specifications, for example, using the patch generation module 311. Color specification conversions can be selected based on the desired type of image enhancement (e.g., emphasizing or enhancing a particular color channel, saturation level, brightness level, etc.). Color specification conversions can also be selected to improve compatibility between the digital pathology image generation system 320 and the digital pathology image processing system 310. For example, certain image scanning components can provide output in HSL color specifications, and as described herein, models used in the digital pathology image processing system 310 can be trained using RGB images. Converting patches to compatible color specifications can ensure that the patches can still be analyzed.Additionally, the digital pathology imaging system can upsample or downsample images provided at a particular color depth (e.g., 8-bit, 16-bit, etc.) so that they can be used by the digital pathology imaging system 310. Furthermore, the digital pathology imaging system 310 can transform the patches depending on the type of image captured (e.g., a fluorescence image may contain more detailed color intensities or a wider range of colors).
[0059] As described herein, the patch classification module 312 can identify or classify image features within patches of the digital pathology image and generate labels for these features. In some embodiments, classifying image features (e.g., features in the digital pathology image) can include classifying or identifying tissue structures underlying a tissue sample. The patch classification module 312 can receive a set of patches from the patch generation module 311, identify one or more features within each patch, and generate one or more labels for these features using a machine learning model. Each label can indicate a particular type of condition (e.g., histological subtype, variant) exhibited in the tissue sample. As an example, the digital pathology image can be an image of a sample from a patient diagnosed with a type of non-small lung cancer, and the features identified by the patch classification module 312 can include different histologies such as adenocarcinoma (ADC), squamous cell carcinoma (SCC), etc., as shown in FIG. 1A. As another example, the features identified by patch classification module 312 can include different mutations or genetic alterations, such as KRAS mutations, epidermal growth factor receptor (EGFR) mutations, anaplastic lymphoma kinase (ALK) mutations, or tumor protein 53 (TP53) mutations, as shown in FIG. 2A. In certain embodiments, patch classification module 312 can identify image features and generate corresponding labels for application to patches using a trained machine learning model, such as deep learning neural network 125 or 225 described herein. The model can be trained by training controller 317 based on the process described below with reference to FIGS. 5A-5B or by the method described in FIG. 8.
[0060] As described herein, the image visualization module 313 can generate visualizations for analyzing digital pathology images. In certain embodiments, the image visualization module 313 can generate a visualization of a given digital pathology image based on features identified in the image, labels corresponding to tissue structure features and generated for patches of the digital pathology image, and other relevant information. For example, the image visualization module 313 can receive labels or labeled patches from the patch classification module 312 and generate a visualization based on the labeled patches, as described in FIGS. 1B and 2B, for example. In certain embodiments, one or more visualizations, such as those shown in FIGS. 4A-4B, can be used by a pathologist to evaluate a machine learning model described herein. For example, a pathologist can review a visualization (e.g., a heat map) showing various identified features and corresponding labels to evaluate whether the model correctly identifies these features and labels. Additionally or alternatively, the one or more visualizations can assist a pathologist in diagnosing or evaluating a patient or reviewing an initial assessment.
[0061] In certain embodiments, the visualization generated by the image visualization module 313 is a patch-based signature, such as a heat map, that characterizes the details of the identified features for review and / or analysis. Note that a heat map is just one type of patch-based signature; other types of patch-based signatures can also be generated and used in the visualizations described herein. In some embodiments, the digital pathology image processing system 310 can learn the patch-based signature and use that learning for other predictions. This can include, for example, visualization of raw counts, percentage of labeled patches, percentage of labeled patches relative to the rest of the slide / tumor area, statistical distribution of labeled patches, spatial distribution of patches, etc.
[0062] The patch-based signature can represent a visualization of the identified features in the tissue sample. For example, the visualization can include displaying features (e.g., histology, mutations) with different color coding, as described elsewhere herein. In certain embodiments, the image visualization module 313 can use different visualization techniques to generate its visualization (e.g., patch-based signature). For example, the image visualization module 313 can generate the visualization using one or more of a gradient-weighted class activation mapping (Grad-CAM) technique, a score-weighted class activation mapping (score-CAM) technique, an occlusion mapping technique, and a saliency mapping technique, as shown and described in FIG. 4C , for example. In one embodiment, the image visualization module 313 can select a saliency mapping technique as a preferred and desired technique for visualizing various features in the image. FIG. 4D illustrates an example of visualizing an image with a selected saliency mapping technique.
[0063] As described herein, the heterogeneity metric calculation module 314 can calculate a heterogeneity metric based on features and / or labels identified in the digital pathology image. The heterogeneity metric can include a quantifiable measure of the level or degree of heterogeneity of features, including histology, mutations, etc., based on the labels corresponding to those features (e.g., histology subtype, variant type). In certain embodiments, the heterogeneity metric can use the labels to indicate the relative proportion of each feature in the tissue structure relative to other features in the tissue structure. By way of example, the heterogeneity metric can include raw counts of labels, percentage of labeled patches, percentage of labeled patches relative to the rest of the slide and / or tumor area, statistical distribution of labeled patches, spatial distribution of labeled patches, and other related metrics and their derivations.
[0064] By way of example, and not limitation, a heterogeneity metric can indicate the respective proportions of ADC and SCC subtypes of lung cancer in a tissue sample (e.g., a patient's non-small cell lung cancer histology) for the various histologies identified in FIG. 1A. In other words, the heterogeneity metric can indicate how many total ADC cancer areas are present relative to the total SCC cancer areas in a tissue sample, as shown in FIG. 1B. As another example, a heterogeneity metric can indicate the respective proportions of KRAS and EGFR mutations in a given tissue sample of a subject or patient for the various mutations or genetic variants identified in FIG. 2A. The heterogeneity metric can be used to diagnose specific conditions in patients. For example, the heterogeneity metric can be used to diagnose a patient's heterogeneity based on the amount of ADC and SCC cancer areas present in the patient's tissue sample.
[0065] The output generation module 316 of the digital pathology image processing system 310 can generate output corresponding to the digital pathology image received as input using the digital pathology image, image classification (e.g., label patches), image visualization (e.g., patch-based signatures), and heterogeneity metrics. In addition to labels and annotations of the digital pathology image, as described herein, the output can include various visualizations and diagnoses corresponding to these visualizations. The output can further include a subject assessment based on a tissue sample. For example, the output for a given digital pathology image can include a so-called heat map that identifies and highlights regions of interest within the digital pathology image, as shown in Figures 4A-4B. The heat map can indicate portions of the image that display or correlate with a particular symptom or diagnosis, and can indicate the accuracy or statistical confidence of such an indication. In many embodiments, the output is provided to a user device 330 for display, but in certain embodiments, the output can be accessed directly from the digital pathology image processing system 310.
[0066] The training controller 317 of the digital pathology imaging system 310 can control the training of one or more machine learning models (e.g., deep learning neural networks) described herein and / or functions used by the digital pathology imaging system 310. In some cases, one or more of the neural networks used by the digital pathology imaging system 310 that are used to identify or detect features (e.g., histology, mutations, etc.) within tissue samples are trained together by the training controller 317. In some cases, the training controller 317 can selectively train models for use by the digital pathology imaging system 310. For example, the digital pathology imaging system 210 can use a first training technique to train a first model for feature classification in digital pathology images, a second training technique to train a second model for calculating a heterogeneity metric, and a third training technique to train a third model for identifying tumor regions or areas in digital pathology images. Training a machine learning model (eg, a deep learning neural network) is described in detail below with reference to at least processes 500 and 550 of FIGS. 5A-5B and method 800 of FIG.
[0067] FIG. 4A illustrates an exemplary visualization that can be generated based on image features and corresponding labels identified using the machine learning models described herein. In particular, FIG. 4A illustrates an exemplary heatmap 400 and a detailed view 410 of the same heatmap 400. A heatmap may be composed of multiple cells. In certain embodiments, the heatmap may be an example of a patch-based signature. The cells of the heatmap may directly correspond to patches generated from the digital pathology image, as described above with reference to at least FIGS. 1A-1B and 2A-2B. Each cell may be assigned an intensity value, which may be normalized across all cells (e.g., so that the intensity values of a cell range from 0 to 1, 0 to 100, etc.). When displaying the heatmap 400, the intensity values of the cells may be translated into different colors, patterns, intensities, etc. In the example shown in FIG. 4A, 402 represents a tumor region or tumor areas within tumor region 402, where dark gray regions 404a, 404b, 404c, 404d, and 404e (individually and collectively also referred to herein as 404) represent ADC cancer regions, and light gray cells 406a, 406b, 406c, and 406d (individually and collectively also referred to herein as 406) represent SCC cancer regions. Although not shown, different mutations (e.g., KRAS mutations, EGFR mutations, etc.) can be similarly visualized using heat maps as discussed herein. Color gradients can be used to illustrate different histologies or mutations, as identified using the processes discussed in FIGS. 1A-1B and 2A-2B. In certain embodiments, the intensity value of each cell can be derived from a label determined for the corresponding patch by a deep learning neural network as described herein. Thus, the heat map is used to enable the digital pathology image processing system 310, and in particular the patch classification module 312, to quickly identify patches of the digital pathology image that have been identified as likely to contain indicators of a particular condition, such as a particular histology, a particular genetic mutation, or a particular genetic alteration.In certain embodiments, visualizations 400 and 410 shown in FIG. 4A are generated using one or more of visualization tool 160, visualization tool 260, image visualization module 313, or output generation module 316.
[0068] Figure 4B shows another example visualization that can be generated based on image features and corresponding labels identified using the machine learning model described herein. In particular, Figure 4B illustrates an example in which two heatmaps 420 and 430 can be generated for a single digital pathology image. Each heatmap 420 or 430 represents a visualization for a single label indicating a particular feature. For example, rather than a single heatmap representing a visualization of multiple features, such as heatmap 400, a separate heatmap can be generated for each feature, with each heatmap for a particular feature representing that feature's region within the map. As shown, heatmap 420 represents SCC regions 422a, 422b, 422c, and 422d, while heatmap 430 represents ADC regions 432a, 432b, and 432c. Similarly, separate heatmaps can be generated to visualize each mutation or genetic variant.
[0069] 4C illustrates the visualization of an exemplary digital pathology image using various visualization techniques, including gradient-weighted class activation mapping (Grad-CAM), score-weighted class activation mapping (score-CAM), occlusion mapping, and saliency mapping. In certain embodiments, the image visualization module 313 of the digital pathology image processing system 310 can generate various patch-based signatures using each of these different visualization techniques.
[0070] As shown, image 450 shows the original patch before applying the visualization technique. Image 452 shows the patch after applying the Grad-CAM technique. The Grad-CAM technique uses the gradient of any target concept fed into the final convolutional layer of a convolutional neural network (CNN) to generate a coarse localization map that highlights important regions in the image for concept prediction. Next, image 454 shows the patch after applying the Score-CAM technique. The Score-CAM technique is a gradient-free visualization method extended from Grad-CAM and Grad-CAM++. It achieves better visual performance and fairness for interpreting decision-making processes. Next, image 456 shows the patch after applying the occlusion mapping technique. The occlusion mapping technique is a shadowing technique used to make 3D objects appear more realistic by simulating the soft shadow that would naturally occur when indirect or ambient light is projected onto an image. In some embodiments, the occlusion map is a grayscale image, with white indicating areas that should receive full indirect light and black indicating no indirect light. Next, image 458 shows the patch after applying a saliency mapping technique. Saliency mapping is a technique that uses saliency to identify unique features (pixels, resolution, etc.) within an image. The unique features indicate important or relevant locations within the image. In particular embodiments, the saliency mapping technique identifies regions within the image that a machine learning model (e.g., a deep learning neural network) uses to make its label predictions. In particular embodiments, the saliency map is also a heat map, where hotness refers to regions of the image that have a significant impact on predicting the class to which an object belongs. The goal of a saliency map is to find salient or prominent regions at all locations within the field of view and guide the selection of attention locations based on the spatial distribution of saliency.
[0071] As described above with reference to FIG. 4C , based on comparing different visualization techniques and the results obtained based on these techniques, the image visualization module 313 can select a preferred technique for visualizing various features (e.g., mutations) within the image. Furthermore, a user operator of the image visualization module 313 can specify the desired technique. For example, based on analyzing or inspecting patches 450, 452, 454, 456, and 458 obtained after applying different visualization techniques in FIG. 4C , the image visualization module 313 can determine the saliency map results as being the most accurate and clearest in indicating features (e.g., ADC, SCC) within the image. FIG. 4D illustrates an example of visualizing an exemplary digital pathology image using a selected saliency mapping technique. Similar to FIG. 4C , 460 illustrates the original SCC patch before applying any visualization technique, and 462 illustrates the SCC patch obtained after applying the saliency mapping technique described herein. The application of the saliency mapping technique helps indicate the most affected areas of tissue in the deep learning neural network's decision to make a prediction regarding the label. As indicated by reference numeral 470, the saliency map extracts the cell nuclei within the SCC patch without a segmentation algorithm.
[0072] 5A-5B illustrate exemplary processes 500 and 550 for classifying digital pathology images according to the detection of image features depicted in patches and for training a machine learning model for testing and updating the machine learning model, respectively. FIG. 5A illustrates an exemplary process 500 for training the digital pathology image processing system 310, particularly for training a deep learning neural network used to identify features (e.g., histology, mutations, etc.) within tissue samples. Generally, the training process involves providing training data (e.g., digital pathology images or whole slide images of various subjects) having ground truth features and corresponding labels to the digital pathology image processing system 310, and teaching the deep learning neural network to identify various features (e.g., histology, mutations, etc.) and generate corresponding labels within a given digital pathology or whole slide image. The ground truth labels used for training may be provided by a clinician / pathologist and may include, for example, tumor type diagnoses, such as histological subtypes including adenocarcinoma and squamous cell carcinoma, and DNA sequencing of mutations. Training is particularly advantageous because it reduces the burden on users (e.g., pathologists, physicians, clinical specialists, etc.) of manually evaluating thousands of entire slide images and identifying features in each of these images. Using a trained machine learning model to identify histology or mutations within tissue samples expedites the overall image classification and evaluation process. Once sufficiently trained, the model can reduce the number of errors or opportunities for error that can be introduced into manual classification and evaluation by humans. For example, a trained model may be able to automatically identify different mutations or genetic variants that would be difficult or impossible to predict manually by a human from H&E images alone. Trained machine learning models may also help identify novel biomarkers or features that were previously unknown.
[0073] In some embodiments, this type of learning structure model can be referred to as multi-instance learning. In multi-instance learning, a collection of instances is provided together as a labeled set. Note that individual instances are often not labeled solely in the set. The labels are typically based on the symptoms present. A fundamental assumption in the multi-instance learning technique employed by the described system is that if a set of patches is labeled as having a condition present (e.g., if a set of patches is labeled as being associated with a particular mutation type), then at least one instance in the set is of the particular mutation type. Similarly, if a set of patches is labeled as being associated with a particular histology, then at least one instance in the set is of the particular histology. In other embodiments, the patches may be individually labeled, and a set of patches may include individually labeled patches, where the label associated with one patch in the set is different from the label associated with another patch in the set.
[0074] As described herein, the training controller 317 of the digital pathology imaging system 310 can control the training of one or more machine learning models (e.g., deep learning neural networks) described herein and / or the functions used by the digital pathology imaging system 310 to identify features (e.g., histology, mutations, etc.) within tissue samples. As shown in FIG. 5A , at 510, the training controller 317 can select, acquire, and / or access training data including a set of digital pathology images (e.g., full slide images 505 a, 505 b, ..., 505 n). Note that while three images are shown in FIG. 5A as part of the training data, this is by no means limiting and any number of images can be used in training. For example, 1000 digital pathology images including tissue samples from NSCLC patients can be used to train the machine learning models (e.g., deep learning neural networks) discussed herein.
[0075] At 520, the training controller 317 can perform tumor lesion segmentation, e.g., identify tumor regions (e.g., disease regions) in each of the digital pathology images. For example, as shown in FIG. 5A , tumor region 515a is identified in image 505a, tumor region 515b is identified in image 505b, and tumor region 515n is identified in image 505n. In one embodiment, the training controller 317 automatically identifies tumor regions using one or more machine learning techniques. For example, a machine learning model can be trained based on pre-labeled or pre-annotated tumor regions in a set of digital pathology images by a human expert. In some embodiments, the training controller 317 can identify tumor regions with the assistance of a user. For example, tumor regions may be manually selected by a pathologist, physician, clinical specialist, lung cancer diagnostic expert, etc.
[0076] At 530, the training controller 317, using, for example, the patch generation module 311, causes the digital pathology image processing system 310 to subdivide each digital pathology image having an identified tumor region into a set of patches or tiles. For example, as shown in FIG. 5A , image 525a is subdivided into a set of patches 535a, 535b, ..., 535n (also individually and collectively referred to herein as 535), image 525b is subdivided into a set of patches 536a, 536b, ..., 536n (also individually and collectively referred to herein as 536), and image 525n is subdivided into a set of patches 537a, 537b, ..., 537n (also individually and collectively referred to herein as 537). Each patch in sets 535, 536, and 537 can be classified by one or more image features and annotated with corresponding labels of those features. For example, one or more human experts or pathologists can annotate each patch with a label that identifies or indicates a specific feature within the tissue sample. By way of example and not limitation, a pathologist can classify or identify one or more features within each patch and provide a label for each identified feature, with the label indicating a specific histology, such as ADC or SCC, within the tissue sample. As another example, but not limited to, a label indicating a specific mutation or genetic variant, such as KRAS, ALK, or TP53, within the tissue sample can be provided by the pathologist for each patch. This process can be repeated until all extracted patches have been classified, annotated, or labeled. In certain embodiments, a single ground truth label for a patient can be used to label each patch (belonging to a tumor lesion) of an image of a sample from the patient. In practice, the labels of patches may be different and not the same, and thus training is weakly supervised.
[0077] It should be understood that the training process 500 shown in FIG. 5A is not limited to the order or arrangement of steps 510, 520, 530, and 540 shown in FIG. 5A , and that rearrangement of one or more steps of the training process 500 is possible and within the scope of the present disclosure. For example, in one embodiment, the subdivision or tiling step 530 can occur before the tumor lesion segmentation step 520. In this embodiment, the subject's digital pathology image 505 a, 505 b, ..., 505 n is first subdivided into multiple patches, and then tumor lesion segmentation is performed on each of these patches (e.g., tumor regions are identified in each patch) using either manual annotation or a separate tumor segmentation algorithm. In another embodiment, the training process 500 is performed in the order shown in FIG. 5A . Other variations and one or more additional steps in the training process 500 are possible and contemplated.
[0078] In step 540, the training controller 317 can train a machine learning model (e.g., a deep learning neural network) based on the set of labeled patches 535, 536, and 537. For example, the training controller 317 can feed each labeled patch (e.g., a patch having identified features and corresponding labels) to the machine learning model for training using CNN training techniques understood by those skilled in the art. Once trained, the machine learning model can classify the tissue patches using whole-slide-level labels, as described elsewhere herein.
[0079] FIG. 5B illustrates a process 550 for testing and updating a machine learning model trained to identify and classify features within and / or patches segmented from digital pathology images using process 500 of FIG. 5A. For example, once a machine learning model has been trained based on multiple digital pathology images 505a, 505b, ..., 505n and corresponding patch labels 535a, 535b, ..., 535n, 536a, 536b, ..., 536n, and 537a, 537b, ..., 537n, as described above with reference to FIG. 5A, the trained machine learning model may be tested on one or more unseen test slides or digital pathology images to verify the accuracy of the trained machine learning model in its classification. As an example, the trained machine learning model may be configured to perform its testing on 20 unseen test slides for validation. The number of test slides or images for testing the machine learning model may be any number and may be preset by a user. Based on the validation, the reliability of the model may be determined.
[0080] At 560, the training controller 317 can access a particular digital pathology image 565 of a particular subject to test the trained machine learning model. At step 570, the digital image processing system 310 can subdivide the particular digital pathology image 565 into a plurality of patches 575a, 575b, ..., 575n (individually and collectively referred to herein as 575). At 580, the training controller 317 uses the trained machine learning model obtained by process 500 of FIG. 5A to identify image features and generate labels for the identified image features in the plurality of patches 575. For example, the training controller 317 generates predicted label 585a for the feature shown in patch 575a, predicted label 585b for the feature shown in patch 575b, and predicted label 585n for the feature shown in patch 575n.
[0081] At 590, the training controller 317 can access the ground truth labels or classifications for each of the patches 575a, 575b, ..., 575n. As shown, ground truth label 587a corresponds to the feature depicted in patch 575a, ground truth label 587b corresponds to the feature depicted in patch 575b, and ground truth label 587n corresponds to the feature depicted in patch 575n. In particular embodiments, the ground truth labels are labels or classifications that are known to be accurate or ideal classifications. For example, the ground truth labels can be provided as part of a dataset of training images or can be generated by a pathologist or other human operator. Upon accessing the ground truth labels, at step 590, the training controller 317 can compare the predicted labels 585a, 585b, ..., 585n with the corresponding ground truth or true labels 587a, 587b, ..., 587n. For example, training controller 317 compares predicted label 585a to ground truth label 587a, predicted label 585b to ground truth label 587b, and predicted label 585n to ground truth label 587n. In some embodiments, based on the comparisons, training controller 317 can calculate a scoring function of the training process, such as a loss function. The scoring function (e.g., loss function) can quantify the discrepancy in classification between the predicted label by the deep learning neural network and the ground truth label. For example, the loss function can indicate an offset value that describes how far the predicted label by the machine learning model is from the ground truth or true label. A comparison of the predicted labels to the true labels is shown, for example, in FIG. 6.
[0082] Based on the comparison 590, the training controller 317 can decide whether to stop training or update the machine learning model (e.g., a deep learning neural network). For example, the training controller 317 can decide to train the deep learning neural network until the loss function indicates that the deep learning neural network has exceeded a threshold of agreement between the predicted labels 585a, 585b, ..., 585n and the ground truth labels 587a, 587b, ..., 587n. In some embodiments, the training controller 317 can decide to train the deep learning neural network for a set number of iterations or epochs. For example, the deep learning neural network can be repeatedly trained and updated using the same set of labeled patches 535, 536, 537 until a specified number of iterations is reached or until some threshold criteria is met. The training controller 317 can also perform multiple iterations to train the deep learning neural network using different training images. The deep learning neural network can also be validated using a reserved test set of images. In some embodiments, the training controller 317 can periodically pause training and provide a test set of patches for which the appropriate labels are known. The training controller 317 can evaluate the output of the deep learning neural network against the known labels on the test set to determine the accuracy of the deep learning neural network. When the accuracy reaches a set threshold, the training controller 317 can stop training the deep learning neural network.
[0083] In some embodiments, once the training controller 317 determines that training is complete, the training controller 317 may output a confidence value indicating the reliability or accuracy of the trained machine learning model (e.g., a deep learning neural network) in its classification. For example, the training controller 317 may output a confidence value of 0.95, indicating that the deep learning neural network is 95% accurate in classifying features in the test images of interest. Examples of confidence values indicating the accuracy of a model are shown, for example, in FIG. 6.
[0084] As described herein, conventional processes for identifying image features and generating corresponding labels for digital pathology images (e.g., entire slide images) are difficult and time-consuming. The digital pathology image processing system 310 and methods for using and training the system described herein can be used to increase the set of images available for training various networks of the digital pathology image processing system. For example, after an initial training pass using data with known labels (potentially including annotations), the digital pathology image processing system 310 can be used to classify patches without existing labels. The generated classifications can be verified by a human agent, and if corrections are necessary, the digital pathology image processing system 310 (e.g., a deep learning neural network) can be retrained using new data. This cycle can be repeated, with human intervention expected to be required for previously unseen examples to improve accuracy rates. Furthermore, once a specified level of accuracy is reached, the labels generated by the digital pathology image processing system 310 can be used as ground truth for training.
[0085] FIG. 6 shows exemplary experimental comparison data between predicted and true labels based on two different sampling data. In particular, FIG. 6 illustrates the accuracy of a machine learning model (e.g., a deep learning neural network) in predicting or identifying histological features such as ADC and SCC for two different numbers of test samples. The chart 600 on the left shows confidence values 610a, 610b, 610c, and 610d, which indicate the model's accuracy in predicting or identifying ADC and SCC regions based on 10 test samples. As shown, the training controller 317 outputs a 0.9 confidence value (indicated by reference numbers 610a and 610d), indicating a 90% match between the true and predicted labels. In other words, the training controller 317 can find that the trained model (e.g., a deep learning neural network) is 90% accurate in correctly identifying ADC and SCC regions within the 10 test samples.
[0086] Chart 620 on the right shows confidence values 630a, 630b, 630c, and 630d indicating the accuracy of the model in predicting or identifying ADC and SCC regions based on 280 test samples. As shown, training controller 317 outputs a 0.76 confidence value (indicated by reference number 630a) in the model identifying ADC in these samples and a 0.92 confidence value (indicated by reference number 630d) in the model identifying SCC in these samples. Specifically, confidence values 630a and 630d indicate a 76% agreement between the true and predicted labels in identifying ADC in these samples and a 92% agreement in identifying SCC in these 280 test samples, respectively.
[0087] FIG. 7 illustrates an exemplary method 700 for identifying features in a digital pathology image using a machine learning model and generating a subject assessment based on the heterogeneity of the identified features in the digital pathology image. Method 700 can begin at step 710, where a digital pathology imaging system 310 receives or otherwise accesses a digital pathology image of a tissue sample. In certain embodiments, the digital pathology image of the tissue sample is a whole-slide image of a sample from a subject or patient diagnosed with non-small cell lung cancer. In one embodiment, the digital pathology image or whole-slide image discussed herein can be an H&E-stained image or obtained from an H&E preparation of a tumor sample from a patient diagnosed with non-small cell lung cancer. As described herein, the digital pathology imaging system 310 can receive images directly from a digital pathology image generation system or can receive images from a user device 330. In other embodiments, the digital pathology imaging system 310 can be communicatively coupled to a database or other system for storing digital pathology images, facilitating the digital pathology imaging system 310 receiving the images for analysis.
[0088] In step 715, the digital pathology image processing system 310 subdivides the image into patches. For example, the digital pathology image processing system 310 may subdivide the image into patches as shown in FIGS. 1A and 2A. As described herein, digital pathology images are expected to be significantly larger than standard images, much larger than is typically feasible for standard image recognition and analysis (e.g., on the order of 100,000 pixels by 100,000 pixels). To facilitate analysis, the digital pathology image processing system 310 subdivides the image into patches. The size and shape of the patches are uniform for analytical purposes, although the size and shape can be variable. In some embodiments, patches can overlap to increase the chance that the image context is properly analyzed by the digital pathology image processing system 310. To balance the effort performed and accuracy, it may be preferable to use non-overlapping patches. Furthermore, subdividing the image into patches can include subdividing the image based on color channels or primary colors associated with the image.
[0089] In step 720, the digital pathology image processing system 310 identifies and classifies one or more image features (e.g., histology, mutations, etc.) within each patch, and in step 825, uses a machine learning model to generate one or more labels for the one or more image features identified within each patch of the digital pathology image, where each label may indicate a particular type of condition within the tissue sample (e.g., cancer type, tumor cell type, mutation type, etc.). In one embodiment, the digital pathology image is an image of a sample from a patient with non-small cell lung cancer, and the labels generated by the machine learning model in step 825 may indicate a tissue subtype such as adenocarcinoma (ADC), squamous cell carcinoma (SCC), etc., as shown in FIG. 1A. In another embodiment, the labels generated by the machine learning model in step 825 may indicate different mutations or gene mutations such as KRAS mutation, epidermal growth factor receptor (EGFR) mutation, anaplastic lymphoma kinase (ALK) mutation, or tumor protein 53 (TP53) mutation, as shown in FIG. 2A. It should be understood that the machine learning models described herein are not limited to generating labels corresponding to different histologies and mutations within a tissue sample, but rather labels corresponding to various other features within a tissue sample can also be generated by the machine learning models. In certain embodiments, the machine learning models described herein are deep learning neural networks.
[0090] In step 730, the digital pathology imaging system 310 can optionally generate a patch-based signature based on the labels generated using the machine learning model described above. For example, the digital pathology imaging system 310 can generate a patch-based signature as shown in FIGS. 4A-4B. In certain embodiments, the patch-based signature is a heat map including multiple regions, each associated with multiple intensity values. One or more of the multiple regions in the heat map may be associated with an indication of a condition in the patient sample, and each intensity value associated with the one or more regions correlates with a statistical confidence of the indication. The patch-based signature can represent a visualization of the identified features in the tissue sample. The visualization can include displaying features (e.g., histology, mutations) with different color coding. By way of example and not limitation, visualizing different histologies in the tissue sample via the patch-based signature can include displaying ADC in blue, SCC in green, etc. In certain embodiments, the digital pathology imaging system 310 can generate the patch-based signatures (e.g., heat maps) described herein using different visualization techniques. For example, the image visualization module 313 may generate the visualization using one or more of the Grad-CAM technique, the Score-CAM technique, the occlusion mapping technique, or the saliency mapping technique, as shown and described in, for example, Figure 4C. In one embodiment, the visualization technique used here is the saliency mapping technique, as shown and described in, for example, Figure 4D.
[0091] In step 735, the digital pathology imaging system 310 calculates a heterogeneity metric using the labels generated in step 725. In an alternative embodiment, the digital pathology imaging system 310 can calculate the heterogeneity metric using the patch-based signatures generated in step 730. In certain embodiments, the heterogeneity metric can indicate the relative proportion of each label in a tissue sample relative to other labels in the tissue sample. By way of example, and not limitation, the heterogeneity metric can indicate the proportion of ADC and SCC cancer regions in a tissue sample (e.g., a patient's non-small cell lung cancer image) for the various histologies identified in FIG. 1A. In other words, the heterogeneity metric can indicate how many total ADC cancer regions there are relative to all SCC cancer regions in a tissue sample, as shown in FIG. 1B. As another example, the heterogeneity metric can indicate the respective proportions of KRAS and EGFR mutations in a given tissue sample of a subject or patient for the various mutations or gene variants identified in FIG. 2A. As another example, the heterogeneity metric may provide or contribute to a metric that quantifies the degree or magnitude of heterogeneity.
[0092] In step 740, the digital pathology image processing system 310 generates a subject assessment based on the calculated heterogeneity or heterogeneity metric. The subject assessment can include, by way of example and not limitation, a subject diagnosis, prognosis, treatment recommendation, or other similar assessment based on the heterogeneity of features in the digital pathology image. For example, based on a heterogeneity metric indicating the degree to which various features (e.g., histology or mutations) and their corresponding labels are heterogeneous in a given tissue sample, the output generation module 316 can generate an appropriate assessment for the given tissue sample. As an example, the assessment can include the severity of a patient's lung cancer based on the amount of ADC and SCC cancer areas present in the patient's tissue sample.
[0093] In step 745, the digital pathology imaging system 310 provides the generated subject assessment to a user, such as a pathologist, physician, clinical specialist, lung cancer diagnostic specialist, or imaging equipment operator. In certain embodiments, the user can use the assessment generated in step 740 to evaluate treatment options for the patient. In some embodiments, the output generation module 316 can output an indication of whether the subject is eligible for a clinical trial based on the assessment. The output (e.g., assessment) can further include, for example, a digital pathology image classification of various image features (e.g., histology, mutations, etc.), an interactive interface, or differential characteristics and statistics therefor. These outputs, etc., can be provided to a user, for example, via a suitably configured user device 330. The output can be provided in an interactive interface that facilitates the user's review of the analysis performed by the digital pathology imaging system 310 while also supporting the user's independent analysis. For example, the user can turn various features of the output on or off, zoom, pan, and otherwise manipulate the digital pathology image, and provide feedback or notes regarding the classification, annotations, and differential characteristics.
[0094] In step 750, the digital pathology imaging system 310 can optionally receive feedback regarding the provided subject assessment. The user can provide feedback regarding the accuracy of the label classification or annotation. The user can, for example, indicate areas of interest to the user (and reasons for interest) that were not previously identified by the digital pathology imaging system 310. The user can further indicate additional classifications of the image that have not yet been suggested or captured by the digital pathology imaging system 310. This feedback can also be stored for later access by the user, for example, as a clinical note.
[0095] In step 755, the digital pathology imaging system 310 can optionally use the feedback to retrain or update one or more of the machine learning models, e.g., deep learning neural networks or classification networks, used to classify the digital pathology images. The digital pathology imaging system 310 can use the feedback to supplement the training dataset available to the digital pathology imaging system 310 with the added advantage that the feedback has been provided by a human expert, increasing its reliability. The digital pathology imaging system 310 can continuously modify the deep learning neural network underlying the analysis provided by the system in order to increase the accuracy of its classification and the speed with which the digital pathology imaging system 310 identifies key regions of interest. Thus, rather than being a static system, the digital pathology imaging system 310 can provide and benefit from continuous improvement.
[0096] Certain embodiments may repeat one or more steps of the method of FIG. 7 , where appropriate. While this disclosure describes and illustrates certain steps of the method of FIG. 7 as occurring in a particular order, this disclosure contemplates any suitable steps of the method of FIG. 7 occurring in any suitable order. Furthermore, while this disclosure describes and illustrates an exemplary method for identifying features in digital pathology images using a machine learning model and generating a subject assessment based on heterogeneity of the identified features in the digital pathology images, including certain steps of the method of FIG. 7 , this disclosure contemplates any suitable method for identifying features in digital pathology images using a machine learning model and generating a subject assessment based on heterogeneity of the identified features in the digital pathology images, including any suitable steps, which may, where appropriate, not include all, some, or any of the steps of the method of FIG. 7 . Furthermore, while this disclosure describes and illustrates certain components, devices, or systems performing certain steps of the method of FIG. 7 , this disclosure contemplates any suitable combination of any suitable components, devices, or systems performing any suitable steps of the method of FIG. 7 .
[0097] FIG. 8 illustrates an exemplary method 800 for training and updating a machine learning model for labeling or classifying patches of digital pathology images according to the detection of image features exhibited in the patches. In certain embodiments, steps 810-830 of method 800 relate to training the machine learning model, and steps 835-865 of method 800 relate to testing and updating the trained machine learning model. Method 800 can begin at step 910, where the digital pathology image processing system 310 accesses a plurality of digital pathology images associated with a plurality of subjects or patients, respectively. In certain embodiments, this includes receiving image data of tissue samples from non-small cell lung cancer (NSCLC) patients. As an example, 476 tissue samples from NSCLC patients scanned at a resolution of 0.5 pixels / μm can be used as a training dataset for training the machine learning model discussed herein.
[0098] In step 815, the digital pathology image processing system 310 performs tumor lesion segmentation, e.g., identifies tumor regions in each of the plurality of digital pathology images accessed in step 810. By way of example, as shown in FIG. 5A , tumor regions 515 can be identified in each image. In one embodiment, tumor regions can be automatically identified using a separate tumor lesion segmentation algorithm or one or more machine learning techniques. For example, a machine learning model can be trained based on pre-labeled or pre-annotated tumor regions in a set of digital pathology images by a human expert. In one embodiment, tumor regions can be manually selected by a user, such as a pathologist, physician, clinical specialist, or lung cancer diagnostic expert.
[0099] In step 820, the digital pathology image processing system 310 may subdivide each digital pathology image having an identified tumor region into a set of patches. For example, as shown in FIG. 5A , the digital pathology image processing system 310 may subdivide each image into a set of patches, such as set 535, set 536, set 537, etc. In some embodiments, the size and shape of each patch is uniform for analysis purposes, although the size and shape may be variable. By way of example, each digital pathology image having an identified tumor region or regions is subdivided into smaller image patches of 512 x 512 pixels. As previously mentioned, this disclosure contemplates that any suitable steps of the training process 500 of FIG. 5A or the method 800 of FIG. 8 may be performed in any suitable order. For example, in one embodiment, step 820 may be performed before step 815. In another embodiment, the steps of the method 800 are performed in the order shown in FIG. 8 .
[0100] In step 825, the set of patches extracted in step 820 may be classified or annotated by image features along with corresponding labels. For example, one or more human experts or pathologists may classify one or more features in each patch and annotate the features with one or more ground truth labels that indicate a particular condition in the tissue sample. By way of example and not limitation, each patch may be classified or labeled as containing a particular histology, such as ADC or SCC, in the tissue sample. By way of another example and not limitation, each patch may be classified or labeled as containing a particular mutation or gene variant, such as KRAS, ALK, or TP53, in the tissue sample. This process may be repeated until all extracted patches have been annotated or labeled.
[0101] In step 830, the digital pathology image processing system 310 can train a machine learning model (e.g., a deep learning neural network) based on the set of labeled patches. For example, the training controller 317 can feed each labeled patch (e.g., classified tissue structure features with corresponding ground truth labels) to the machine learning model for training, as shown, for example, in FIG. 5A. In certain embodiments, the machine learning model is a convolutional neural network that can be trained based on Inception V3 and Resnet18 architectures using transfer learning and weakly supervised learning techniques. It should be understood that other learning techniques for training the machine learning model are also possible and within the scope of the present disclosure. Once trained, the machine learning model can classify tissue patches using whole-slide-level labels.
[0102] In step 835, the digital pathology image processing system 310 can access specific digital pathology images of specific subjects to test the trained machine learning model. For example, once a machine learning model has been trained based on multiple digital pathology images and corresponding patch labels as described above in steps 810-830, the trained machine learning model may be tested on one or more unseen test slides or digital pathology images to verify the accuracy of the trained machine learning model in its classification and determine the reliability of the model. As an example, the trained machine learning model may be configured to perform its testing on 20 unseen test slides for validation. The number of test slides or images for testing the machine learning model may be any number and may be preset by the user.
[0103] In step 840, the digital pathology image processing system 310 may subdivide the particular digital pathology image into a set of second patches, as discussed elsewhere herein and shown, for example, in Figure 5B. In step 845, the digital pathology image processing system 310 may identify one or more second image features within each patch and generate one or more labels (e.g., histology subtype, mutation type, etc.) using the trained machine learning model.
[0104] In step 850, the digital pathology image processing system 310 can compare the labels generated by the trained machine learning model with ground truth or true labels. In some embodiments, the digital pathology image processing system 310 can calculate a loss function based on the comparison. For example, the training controller 317 can compare the predicted labels of the second set of patches by the machine learning model with the true labels of those patches by a human expert or pathologist to determine the loss function. In some embodiments, the loss function can be a measure of the accuracy of the machine learning model in predicting the labels of features represented in a given tissue sample. In some embodiments, the loss function can indicate an offset value that quantifies how far the predicted labels by the machine learning model deviate from the ground truth or true labels. A comparison of the predicted labels to the true labels is shown, for example, in FIG. 6.
[0105] In step 855, the digital pathology image processing system 310 can optionally make a determination as to whether the scoring function (e.g., loss function) calculated based on the comparison in step 850 is below a certain threshold. The threshold may be an upper limit set by a user (e.g., a pathologist) until the labels predicted by the machine learning model for the second set of patches are considered close to or equivalent to the true or ground truth labels. In other words, the threshold may be a limit or value, and if the scoring function indicating an offset value (e.g., quantifying how far the labels predicted by the machine learning model are from the ground truth or true labels) is below or within the threshold, the machine learning model can be determined to be accurate in its label prediction or classification. On the other hand, if the offset value of the scoring function is greater than the threshold, the machine learning model is determined to be inaccurate and is flagged as requiring more training. As a non-limiting example, the threshold may be 90%, and if a comparison between the predicted labels and the true labels reveals a 92% match between the labels, or 92% of the predicted labels match the true labels, the machine learning model can be considered accurate and sufficiently trained. Continuing with the same example, if the match between the predicted labels and the true labels is only 75%, the machine learning model is determined to require or need more training. In some embodiments, the training controller 317 can make this determination in step 855 using the comparison data, for example, as shown in FIG. 6.
[0106] At step 860, the digital pathology image processing system 310 may update the machine learning model. In certain embodiments, the update occurs in response to determining that the scoring function is less than a threshold. In some embodiments, updating the machine learning model may include one or more of repeating steps 810-830, reconfiguring or updating one or more parameters of the machine learning model, and performing steps 835-855 to check whether the loss function meets a threshold criterion (e.g., the loss function is greater than a threshold, the agreement between the predicted labels and the true labels is greater than 90%, etc.). In certain embodiments, the update occurs to optimize the loss function or to minimize the difference between the generated / predicted labels and the true / ground truth labels.
[0107] In step 865, the digital pathology imaging system 310 may terminate training and store the trained machine learning model in a data store for future access and / or retrieval in classifying features (e.g., histology, mutations, etc.) within tissue samples. In some embodiments, the training controller 317 determines when training should be stopped. The determination may be based on predetermined termination rules. In some embodiments, training may terminate in response to determining that the scoring function meets or is greater than a threshold criterion. In certain embodiments, training may terminate when a predetermined number of training samples (e.g., 1,000, 10,000, etc.) have been used to train the model. In certain embodiments, training may terminate when all training samples in the training dataset have been used to train the model. In certain embodiments, training may terminate if the loss comparison (e.g., the offset value of the loss function) is sufficiently small or below a predetermined threshold. If the training controller 317 determines that training should continue, the process may repeat from step 810. Alternatively, if the training controller 317 determines that training should end, then training ends.
[0108] Particular embodiments may repeat one or more steps of the method of FIG. 8 , where appropriate. While this disclosure describes and illustrates certain steps of the method of FIG. 8 as occurring in a particular order, this disclosure contemplates any suitable steps of the method of FIG. 8 occurring in any suitable order. Moreover, while this disclosure describes and illustrates exemplary methods for training and updating machine learning models for labeling or classifying patches of digital pathology images according to detection of image features depicted in the patches that include certain steps of the method of FIG. 8 , this disclosure contemplates any suitable method for training and updating machine learning models for labeling or classifying patches of digital pathology images according to detection of image features depicted in the patches that includes any suitable steps, which may not include all, some, or any of the steps of the method of FIG. 8 , where appropriate. Furthermore, while this disclosure describes and illustrates particular components, devices, or systems that perform certain steps of the method of FIG. 8 , this disclosure contemplates any suitable combination of any suitable components, devices, or systems that perform any suitable steps of the method of FIG. 8 .
[0109] The general techniques described herein can be integrated into a variety of tools and use cases. For example, as described above, a user (e.g., a pathologist or clinician) can access a user device 330 that communicates with the digital pathology imaging system 310 and provide digital pathology images for analysis. The digital pathology imaging system 310, or a connection to the digital pathology imaging system, can be provided as a standalone software tool or package that automatically annotates digital pathology images and / or generates heat maps that evaluate the images under analysis. As standalone tools or plug-ins that can be purchased or licensed on a streamlined basis, the tools can be used to enhance the capabilities of research or clinical laboratories. Additionally, the tools can be integrated into services made available to customers of the digital pathology imaging system. For example, the tools can be provided as a unified workflow, where a user who runs or requests a digital pathology image be created automatically receives the annotated image or heat map equivalent. Thus, in addition to improving digital pathology image analysis, these techniques can be integrated into existing systems to provide additional features not previously considered or possible.
[0110] Additionally, the digital pathology imaging system 310 can be trained and customized for use in a particular setting. For example, the digital pathology imaging system 310 can be specifically trained for use in providing a clinical diagnosis regarding a particular type of tissue (e.g., lung, heart, blood, liver, etc.). As another example, the digital pathology imaging system 310 can be trained to assist in safety assessment, for example, in determining the level or degree of toxicity associated with a drug or other potential therapeutic treatment. Once trained for use in a particular subject or use case, the digital pathology imaging system 310 is not necessarily limited to that use case. For example, the digital pathology imaging system may be trained for use in toxicity assessment of liver tissue, but the resulting model can be applied in a diagnostic setting. Training can be performed in a specific situation, such as toxicity assessment, for a relatively large set of at least partially labeled or annotated digital pathology images.
[0111] 9 illustrates an exemplary computer system 900. In particular embodiments, one or more computer systems 900 perform one or more steps of one or more methods described or illustrated herein. In particular embodiments, one or more computer systems 900 provide functionality described or illustrated herein. In particular embodiments, software executing on one or more computer systems 900 performs one or more steps of one or more methods described or illustrated herein or provides functionality described or illustrated herein. Particular embodiments include one or more portions of one or more computer systems 900. As used herein, references to a computer system can encompass computing devices, and vice versa, where appropriate. Furthermore, references to a computer system can encompass one or more computer systems, where appropriate.
[0112] The present disclosure contemplates any suitable number of computer systems 900. The present disclosure contemplates computer system 900 taking any suitable physical form. By way of example and not limitation, computer system 900 can be an embedded computer system, a system-on-chip (SOC), a single-board computer system (SBC) (e.g., a computer-on-module (COM) or system-on-module (SOM)), a desktop computer system, a laptop or notebook computer system, an interactive kiosk, a mainframe, a mesh of computer systems, a mobile phone, a personal digital assistant (PDA), a server, a tablet computer system, an augmented / virtual reality device, or a combination of two or more of these. Where appropriate, computer system 900 can include one or more computer systems 900, be unitary or distributed, span multiple locations, span multiple machines, span multiple data centers, reside in a cloud that can include one or more cloud components in one or more networks. Where appropriate, one or more computer systems 900 can perform one or more steps of one or more methods described or illustrated herein without substantial spatial or temporal limitations. By way of example, and not limitation, one or more computer systems 900 may perform one or more steps of one or more methods described or illustrated herein in real time or batch mode. One or more computer systems 900 may, where appropriate, perform one or more steps of one or more methods described or illustrated herein at different times or in different locations.
[0113] In a particular embodiment, computer system 900 includes a processor 902, memory 904, storage 906, an input / output (I / O) interface 908, a communication interface 910, and a bus 912. Although this disclosure describes and illustrates a particular computer system having a particular number of particular components in a particular configuration, this disclosure contemplates any suitable computer system having any suitable number of any suitable components in any suitable configuration.
[0114] In particular embodiments, processor 902 includes hardware for executing instructions, such as those comprising a computer program. By way of example and not limitation, to execute instructions, processor 902 may retrieve (or fetch) instructions from an internal register, an internal cache, memory 904, or storage device 906, decode and execute them, and then write one or more results to an internal register, an internal cache, memory 904, or storage device 906. In particular embodiments, processor 902 may include one or more internal caches for data, instructions, or addresses. This disclosure contemplates processor 902 including any suitable number of any suitable internal caches, where appropriate. By way of example and not limitation, processor 902 may include one or more instruction caches, one or more data caches, and one or more translation lookaside buffers (TLBs). Instructions in an instruction cache may be copies of instructions in memory 904 or storage device 906, and the instruction cache may speed up retrieval of those instructions by processor 902. The data in the data cache may be a copy of data in memory 904 or storage 906 for instructions that operate on data, such as the results of a previous instruction executing in processor 902 for access by a subsequent instruction executing in processor 902, or for writing to memory 904 or storage 906, or other suitable data. The data cache may speed up read or write operations by processor 902. The TLB may speed up virtual address translation for processor 902. In particular embodiments, processor 902 may include one or more internal registers for data, instructions, or addresses. This disclosure contemplates processor 902 including any suitable number of any suitable internal registers, where appropriate. Where appropriate, processor 902 may include one or more arithmetic logic units (ALUs), may be a multi-core processor, or may include one or more processors 902.Although this disclosure describes and illustrates a particular processor, this disclosure contemplates any suitable processor.
[0115] In particular embodiments, memory 904 includes a main memory for storing instructions for processor 902 to execute or data for processor 902 to operate on. By way of example and not limitation, computer system 900 may load instructions into memory 904 from storage device 906 or another source (e.g., another computer system 900, etc.). Processor 902 may then load the instructions from memory 904 into an internal register or internal cache. To execute the instructions, processor 902 may retrieve the instructions from the internal register or internal cache and decode them. During or after execution of an instruction, processor 902 may write one or more results (which may be intermediate or final results) to an internal register or internal cache. Processor 902 may then write one or more of those results to memory 904. In particular embodiments, processor 902 executes only instructions in one or more internal registers or caches or memory 904 (as opposed to storage device 906 or elsewhere) and operates only on data in one or more internal registers or caches or memory 904 (as opposed to storage device 906 or elsewhere). One or more memory buses (each of which may include an address bus and a data bus) may couple processor 902 to memory 904. Bus 912, as described below, may include one or more memory buses. In particular embodiments, one or more memory management units (MMUs) reside between processor 902 and memory 904 to facilitate accesses to memory 904 requested by processor 902. In particular embodiments, memory 904 includes random access memory (RAM). This RAM may be volatile memory, where appropriate. Where appropriate, this RAM may be dynamic RAM (DRAM) or static RAM (SRAM). Furthermore, where appropriate, this RAM may be single-ported or multi-ported RAM. This disclosure contemplates any suitable RAM. Memory 904 may include one or more memories 904, where appropriate.Although this disclosure describes and illustrates particular memory, this disclosure contemplates any suitable memory.
[0116] In particular embodiments, storage device 906 includes mass storage for data or instructions. By way of example and not limitation, storage device 906 may include a hard disk drive (HDD), a floppy disk drive, flash memory, an optical disk, a magneto-optical disk, magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more of these. Storage device 906 may include removable or non-removable (i.e., fixed) media, where appropriate. Storage device 906 may be internal or external to computer system 900, where appropriate. In particular embodiments, storage device 906 is non-volatile solid-state memory. In particular embodiments, storage device 906 includes read-only memory (ROM). Where appropriate, this ROM may be mask-programmed ROM, programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), electrically alterable ROM (EAROM), or flash memory, or a combination of two or more of these. The present disclosure contemplates mass storage device 906 taking any suitable physical form. Storage 906 may include, where appropriate, one or more storage control units that facilitate communication between processor 902 and storage 906. Where appropriate, storage 906 may include one or more memory storage devices 906. Although this disclosure describes and illustrates particular storage devices, this disclosure contemplates any suitable storage device.
[0117] In particular embodiments, I / O interface 908 includes hardware, software, or both that provide one or more interfaces for communication between computer system 900 and one or more I / O devices. Computer system 900 may include one or more of these I / O devices, where appropriate. One or more of these I / O devices may enable communication between a person and computer system 900. By way of example and not limitation, an I / O device may include a keyboard, keypad, microphone, monitor, mouse, printer, scanner, speaker, still camera, stylus, tablet, touch screen, trackball, video camera, another suitable I / O device, or a combination of two or more of these. An I / O device may include one or more sensors. This disclosure contemplates any suitable I / O devices and any suitable I / O interface 908 therefor. Where appropriate, I / O interface 908 may include one or more device or software drivers that enable processor 902 to drive one or more of these I / O devices. I / O interface 908 may, where appropriate, include one or more I / O interfaces 908. Although this disclosure describes and illustrates particular I / O interfaces, this disclosure contemplates any suitable I / O interface.
[0118] In particular embodiments, communication interface 910 includes hardware, software, or both that provide one or more interfaces for communications (e.g., packet-based communications, etc.) between computer system 900 and one or more other computer systems 900 or one or more networks. By way of example and not limitation, communication interface 910 may include a network interface controller (NIC) or network adapter for communicating with an Ethernet or other wired-based network, or a wireless NIC (WNIC) or wireless adapter for communicating with a wireless network, such as a Wi-Fi network. This disclosure contemplates any suitable network and any suitable communication interface 910 therefor. By way of example and not limitation, computer system 900 may communicate with an ad hoc network, a personal area network (PAN), a local area network (LAN), a wide area network (WAN), a metropolitan area network (MAN), or one or more portions of the Internet, or a combination of two or more of these. One or more portions of one or more of these networks may be wired or wireless. By way of example, computer system 900 may communicate with a wireless PAN (WPAN) (e.g., a BLUETOOTH WPAN, etc.), a Wi-Fi network, a Wi-MAX network, a cellular network (e.g., a Global System for Mobile Communications (GSM) network), or any other suitable wireless network, or a combination of two or more thereof. Computer system 900 may include any suitable communication interface 910 for any of these networks, where appropriate. Communication interface 910 may include one or more communication interfaces 910, where appropriate. Although this disclosure describes and illustrates a particular communication interface, this disclosure contemplates any suitable communication interface.
[0119] In particular embodiments, bus 912 includes hardware, software, or both that couple components of computer system 900 together. By way of example, and not limitation, bus 912 may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an InfiniBand interconnect, a Low Pin Count (LPC) bus, a memory bus, a MicroChannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI Express (PCIe) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or another suitable bus, or a combination of two or more thereof. Bus 912 may include one or more buses 912, where appropriate. Although this disclosure describes and illustrates a particular bus, this disclosure contemplates any suitable bus or interconnect.
[0120] As used herein, a computer-readable non-transitory storage medium may include one or more semiconductor-based or other integrated circuits (ICs) (such as, for example, field programmable gate arrays (FPGAs) or application-specific ICs (ASICs)), hard disk drives (HDDs), hybrid hard drives (HHDs), optical disks, optical disk drives (ODDs), magneto-optical disks, magneto-optical drives, floppy disks, floppy disk drives (FDDs), magnetic tapes, solid-state drives (SSDs), RAM drives, secure digital cards or drives, any other suitable computer-readable non-transitory storage medium, or any suitable combination of two or more of these, where appropriate. A computer-readable non-transitory storage medium may be volatile, non-volatile, or a combination of volatile and non-volatile, where appropriate.
[0121] As used herein, "or" is inclusive and not exclusive, unless expressly indicated otherwise or indicated otherwise by context. Thus, as used herein, "A or B" means "A, B, or both," unless expressly indicated otherwise or indicated otherwise by context. Furthermore, "and" is both conjunctive and several, unless expressly indicated otherwise or indicated otherwise by context. Thus, as used herein, "A and B" means "A and B, together or separately," unless expressly indicated otherwise or indicated otherwise by context.
[0122] The scope of the present disclosure encompasses all modifications, substitutions, variations, alterations, and alterations to the exemplary embodiments described or illustrated herein that would be understood by a person skilled in the art. The scope of the present disclosure is not limited to the exemplary embodiments described or illustrated herein. Furthermore, although the present disclosure describes and illustrates each embodiment herein as including particular components, elements, features, functions, operations, or steps, any of these embodiments may include any combination or permutation of any of the components, elements, features, functions, operations, or steps described or illustrated anywhere herein that would be understood by a person skilled in the art. Furthermore, a reference in the appended claims to an apparatus or system or a component of a system that is adapted, arranged, enabled, configured, enabled, operative, or operates to perform a particular function encompasses that apparatus, system, or component, so long as it is so adapted, arranged, enabled, configured, enabled, operative, or operates, regardless of whether it or that particular function is activated, turned on, or unlocked. Furthermore, while this disclosure describes or illustrates particular embodiments as providing certain advantages, a particular embodiment may provide none, some, or all of these advantages.
Claims
1. 1. A computer-implemented method comprising: receiving a digital pathology image of the tissue sample; subdividing the digital pathology image into a plurality of patches; For each patch of the plurality of patches, identifying image features detected in the patch; generating one or more labels corresponding to the image features identified in the patch using a trained machine learning model; and determining a heterogeneity metric for the tissue sample based on the generated labels; generating a patch-based signature indicative of the degree of heterogeneity of the identified image features and corresponding labels; generating an assessment of the tissue sample based on the heterogeneity metric; A method comprising:
2. 10. The method of claim 1, wherein the digital pathology image of the tissue sample is a whole slide image of a tumor sample from a patient diagnosed with non-small cell lung cancer.
3. The method of claim 1 , wherein the digital pathology image is a hematoxylin and eosin (H&E) stained image.
4. The method of claim 1 , wherein the image features detected in the plurality of patches of the digital pathology image correspond to histological features.
5. The method of claim 4 , wherein the generated labels correspond to cancer regions of adenocarcinoma and squamous cell carcinoma.
6. 10. The method of claim 1, wherein image features detected in the plurality of patches of the digital pathology image correspond to mutations or genetic variants.
7. 7. The method of claim 6, wherein the generated label corresponds to a Kirsten rat sarcoma viral oncogene homolog (KRAS) mutation, an epidermal growth factor receptor (EGFR) mutation, an anaplastic lymphoma kinase (ALK) mutation, or a tumor protein 53 (TP53) mutation.
8. The method of claim 1 , wherein the heterogeneity metric indicates the degree of heterogeneity of the identified image features and corresponding labels in the tissue sample.
9. Using the patch-based signatures to generate a visualization of image features or to evaluate the machine learning model. The method of claim 1 further comprising:
10. 10. The method of claim 9, wherein the patch-based signature comprises a heat map, the heat map comprising a plurality of regions, each region of the plurality of regions being associated with a respective intensity value, and one or more of the plurality of regions being further associated with a predicted label of the patch of the digital pathology image.
11. 10. The method of claim 9, wherein the visualization of the image features identified in the tissue sample comprises displaying each of the generated labels in a distinct color.
12. The method of claim 9 , wherein the patch-based signatures are generated using a saliency mapping technique.
13. further comprising training the machine learning model, wherein training the machine learning model comprises: accessing a plurality of digital pathology images respectively associated with a plurality of subjects; identifying a tumor region in each of the plurality of digital pathology images; subdividing each of the plurality of digital pathology images into a set of training patches, each training patch in the set being classified by one or more features and annotated with one or more ground truth labels corresponding to the one or more features; training the machine learning model using a classified set of training patches having ground truth labels corresponding to the features depicted in the patches; The method of claim 1 , comprising:
14. The method of claim 13 , wherein the ground truth labels are provided by a clinician.
15. further comprising updating the machine learning model, wherein updating the machine learning model comprises: accessing a particular digital pathology image of a particular subject; subdividing the particular digital pathology image into a second set of patches; identifying a second image feature within the second set of patches; using the trained machine learning model to generate a set of predicted labels corresponding to the second image features identified in the second set of patches; comparing the generated set of predicted labels with the ground truth labels to measure the accuracy of label predictions of the machine learning model; updating the machine learning model based on the comparison, where updating the machine learning model comprises further training the machine learning model; and 14. The method of claim 13, comprising:
16. The method of claim 1 , wherein the machine learning model is a deep learning neural network.
17. 10. The method of claim 1, further comprising determining one or more treatment options for the patient based on the evaluation.
18. 1. A digital pathology image processing system, comprising: one or more processors; one or more computer-readable non-transitory storage media coupled to one or more of the processors and comprising instructions operable, when executed by one or more of the processors, to cause the system to perform operations, the operations including: receiving a digital pathology image of the tissue sample; subdividing the digital pathology image into a plurality of patches; For each patch of the plurality of patches, identifying image features detected in the patch; generating one or more labels corresponding to the image features identified in the patch using a machine learning model; and determining a heterogeneity metric for the tissue sample based on the generated labels; generating a patch-based signature indicative of the degree of heterogeneity of the identified image features and corresponding labels; generating an assessment of the tissue sample based on the heterogeneity metric; A digital pathology imaging system, comprising:
19. The digital pathology image processing system of claim 18 , wherein the image features detected in the plurality of patches of the digital pathology image correspond to histological images or mutations.
20. One or more computer-readable non-transitory storage media containing instructions that, when executed by one or more processors, cause the one or more processors of a digital pathology imaging system to: receiving a digital pathology image of the tissue sample; subdividing the digital pathology image into a plurality of patches; For each patch of the plurality of patches, identifying image features detected in the patch; generating one or more labels corresponding to the image features identified in the patch using a machine learning model; and determining a heterogeneity metric for the tissue sample based on the generated labels; generating a patch-based signature indicative of the degree of heterogeneity of the identified image features and corresponding labels; generating an assessment of the tissue sample based on the heterogeneity metric; 1. A computer-readable non-transitory storage medium configured to cause a computer to perform operations including:
Citation Information
Patent Citations
Classification and mutation prediction from histopathology images using deep learning
US20200184643A1
Detecting intratumor heterogeneity of molecular subtypes in pathology slide images using deep-learning
WO2019108695A1